readable AI benchmarks (simplified)

readable AI benchmarks

Alibaba

Qwen3.8 2.4T A95B

reasoning · open weights · Aug 12, 2026

Qwen3.8 2.4T A95B has a Quality Score of 56.8, ranking 72nd among 571 scored models. Per 1M tokens, pricing is $2.00 input, $6.00 output, and $0.250 cached input. It accepts text input, outputs text, and has a 262k token context window. Its Reliability Score is 52.2, ranking 105th among 573 scored models, above average (total average: 36.5). Its Value Score is 38.9, high compared with the total Value average of 31.4.

Quality Score56.8coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score38.9
Factual reliability52.2
Cache Discount88%

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
262k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$2.00
Output
$6.00
Cached Input
$0.250
Cache write
not available
Cost per task
$2.16

Capability

Intelligence
39.9
Coding
not available
Agentic
not available
Omniscience
4.3
Correct
31.3%
Blended price
$1.18

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 58.2%

Correct 31.3%Incorrect 26.9%Partial / not attempted 41.8%

Artificial Analysis benchmarks

GDPval-AA v2
56%
τ³-Banking
49%
SciCode
54%
Humanity’s Last Exam
42%
GPQA Diamond
94%
CritPt
20%
AA-Omniscience
52%
AA-LCR
80%

Similar models

Support me! Patreon