readable AI benchmarks (simplified)

readable AI benchmarks

Anthropic

Claude 4.5 Haiku (Reasoning)

reasoning · closed weights · release date not listed

Claude 4.5 Haiku (Reasoning) has a Quality Score of 36.9, ranking 216th among 571 scored models. Per 1M tokens, pricing is $1.00 input, $5.00 output, $0.100 cached input, and $1.25 cache write. It has a 200k token context window. Its Reliability Score is 47.8, ranking 151st among 573 scored models, above average (total average: 36.5). Its Value Score is 26.6, low compared with the total Value average of 31.4.

Quality Score36.9coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score26.6
Factual reliability47.8
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
200k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$1.00
Output
$5.00
Cached Input
$0.100
Cache write
$1.25
Cost per task
$0.277

Capability

Intelligence
16.9
Coding
not available
Agentic
not available
Omniscience
-4.4
Correct
18.0%
Blended price
$0.770

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 40.4%

Correct 18.0%Incorrect 22.4%Partial / not attempted 59.6%

Artificial Analysis benchmarks

GDPval-AA v2
12%
τ³-Banking
9%
SciCode
42%
Humanity’s Last Exam
10%
GPQA Diamond
67%
CritPt
0%
AA-Omniscience
48%
AA-LCR
74%

Similar models

Support me! Patreon