readable AI benchmarks (simplified)

readable AI benchmarks

Alibaba

Qwen3.7 Max

reasoning · closed weights · May 19, 2026

Qwen3.7 Max has a Quality Score of 54.4, ranking 25th among 53 scored models. Per 1M tokens, pricing is $2.50 input, $7.50 output, $0.250 cached input, and $3.13 cache write. Its Reliability Score is 57.0, ranking 20th among 53 scored models, above average (total average: 50.1). Its Value Score is 36.1, around average compared with the total Value average of 39.6.

Quality Score54.4coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Score36.1
Reliability57.0
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
not available
Weights
closed weights

Token prices USD per 1M tokens

Input
$2.50
Output
$7.50
Cached Input
$0.250
Cache write
$3.13
Cost per task
$1.03

Capability

Intelligence
46.0
Coding
66.0
Agentic
30.6
Omniscience
14.1
Correct
30.1%
Blended price
$1.43

Answer outcomes

Attempt rate: 46.2%

Correct 30.1%Incorrect 16.0%Abstained 53.8%

Artificial Analysis benchmarks

GDPval-AA v2
39%
τ³-Banking
11%
Terminal-Bench v2.1
75%
SciCode
49%
Humanity’s Last Exam
38%
GPQA Diamond
92%
CritPt
13%
AA-Omniscience
57%
AA-LCR
69%

Similar models