readable AI benchmarks

Alibaba

Qwen3.8 2.4T A95B

reasoning · open weights · Aug 12, 2026

Quality Score60.9
Value Score44.3
Reliability52.2
Cache Discount88%

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
983,616 tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$2.00
Output
$6.00
Cached Input
$0.250
Cached Output
not available
Cost per task
$1.09

Capability

Intelligence
57.7
Coding
71.9
Agentic
57.1
Omniscience
4.3
Correct
31.3%
Blended price
$1.18

Answer outcomes

Correct 31.3%Incorrect 26.9%Abstained 41.8%

Artificial Analysis benchmarks

GDPval-AA v2
61%
τ³-Banking
49%
Terminal-Bench v2.1
82%
SciCode
52%
Humanity’s Last Exam
42%
GPQA Diamond
94%
CritPt
20%
AA-Omniscience
52%
AA-LCR
75%

Similar models