readable AI benchmarks (simplified)

readable AI benchmarks

IBM

Granite 4.2 30B

reasoning · open weights · Aug 25, 2026

Granite 4.2 30B has a Quality Score of 23.6, ranking 116th among 265 scored models. Per 1M tokens, pricing is $0.160 input, $0.650 output, and $0.040 cached input. It accepts text input, outputs text, and has a 131,072 token context window. Its Reliability Score is 43.5, ranking 78th among 265 scored models, above average (total average: 33.4). Its Value Score is 22.6, low compared with the total Value average of 27.6.

Quality Score23.6coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Score22.6
Reliability43.5
Cache Discount75%

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
131,072 tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.160
Output
$0.650
Cached Input
$0.040
Cached Output
not available
Cost per task
$0.029

Capability

Intelligence
23.7
Coding
29.9
Agentic
13.9
Omniscience
-12.9
Correct
10.1%
Blended price
$0.125

Answer outcomes

Attempt rate: 33.1%

Correct 10.1%Incorrect 23.0%Abstained 66.9%

Artificial Analysis benchmarks

GDPval-AA v2
14%
τ³-Banking
14%
Terminal-Bench v2.1
27%
SciCode
37%
Humanity’s Last Exam
11%
GPQA Diamond
64%
CritPt
0%
AA-Omniscience
44%
AA-LCR
47%

Similar models