readable AI benchmarks (simplified)

readable AI benchmarks

StepFun

Step 3.7 Flash

reasoning · open weights · May 29, 2026

Step 3.7 Flash has a Quality Score of 30.7, ranking 263rd among 571 scored models. Per 1M tokens, pricing is $0.200 input, $1.15 output, and $0.040 cached input. It accepts text, image, and video input, outputs text, and has a 262,144 token context window. Its Reliability Score is 31.4, ranking 312th among 573 scored models, below average (total average: 36.5). Its Value Score is 26.1, low compared with the total Value average of 31.4.

Quality Score30.7coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score26.1
Factual reliability31.4
Cache Discount80%

Model specification

Reasoning
reasoning
Input modalities
text, image, video
Output modalities
text
Context window
262,144 tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.200
Output
$1.15
Cached Input
$0.040
Cache write
not available
Cost per task
not available

Capability

Intelligence
19.5
Coding
not available
Agentic
not available
Omniscience
-37.3
Correct
25.8%
Blended price
$0.183

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 88.9%

Correct 25.8%Incorrect 63.1%Partial / not attempted 11.1%

Artificial Analysis benchmarks

GDPval-AA v2
18%
τ³-Banking
12%
SciCode
44%
Humanity’s Last Exam
21%
GPQA Diamond
81%
CritPt
2%
AA-Omniscience
31%
AA-LCR
74%

Similar models

Support me! Patreon