readable AI benchmarks (simplified)

readable AI benchmarks

IBM

Granite 4.0 H Small

non-reasoning · open weights · Sep 22, 2025

Granite 4.0 H Small has a Quality Score of 14.4, ranking 513th among 572 scored models. Per 1M tokens, pricing is $0.060 input and $0.250 output. It accepts text input, outputs text, and has a 128k token context window. Its Reliability Score is 19.5, ranking 483rd among 574 scored models, below average (total average: 36.5).

Quality Score14.4coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Scorenot available
Factual reliability19.5
Cache Discountnot available

Model specification

Reasoning
non-reasoning
Input modalities
text
Output modalities
text
Context window
128k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.060
Output
$0.250
Cached Input
not available
Cache write
not available
Cost per task
not available

Capability

Intelligence
6.0
Coding
not available
Agentic
not available
Omniscience
-61.0
Correct
14.3%
Blended price
not available

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 89.6%

Correct 14.3%Incorrect 75.3%Partial / not attempted 10.4%

Artificial Analysis benchmarks

Humanity’s Last Exam
4%
GPQA Diamond
42%
CritPt
0%
AA-Omniscience
19%
AA-LCR
11%

Similar models

Support me! Patreon