readable AI benchmarks (simplified)

readable AI benchmarks

Alibaba

Qwen3.8 Max

reasoning · closed weights · release date not listed

Qwen3.8 Max has a Quality Score of 46.5, ranking 53rd among 296 scored models. Per 1M tokens, pricing is $2.00 input, $6.00 output, and $0.250 cached input. It has a 1M token context window. Its Reliability Score is 51.7, ranking 54th among 296 scored models, above average (total average: 35.7). Its Value Score is 33.4, around average compared with the total Value average of 29.1.

Quality Score46.5coverage 78.8%: missing Epoch General ECI, Coding, CursorBench 3.2, Agentic and DeepSWE v1.1
Value Score33.4
Reliability51.7
Cache Discount88%

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$2.00
Output
$6.00
Cached Input
$0.250
Cache write
not available
Cost per task
$2.67

Capability

Intelligence
40.3
Coding
not available
Agentic
not available
Omniscience
3.4
Correct
31.9%
Blended price
$1.18

Answer outcomes

Attempt rate: 60.3%

Correct 31.9%Incorrect 28.4%Abstained 39.7%

Artificial Analysis benchmarks

GDPval-AA v2
57%
τ³-Banking
51%
SciCode
53%
Humanity’s Last Exam
43%
GPQA Diamond
93%
CritPt
20%
AA-Omniscience
52%
AA-LCR
78%

Similar models