readable AI benchmarks

Google

Gemma 4 31B (Reasoning)

reasoning · open weights · Apr 2, 2026

Quality Score25.6
Value Score33.5
Reliability26.0
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
text, image, video
Output modalities
text
Context window
256k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.0000
Output
$0.0000
Cached Input
not available
Cached Output
not available
Cost per task
$0.0000

Capability

Intelligence
29.7
Coding
43.4
Agentic
14.4
Omniscience
-47.9
Correct
20.0%
Blended price
$0.0000

Answer outcomes

Correct 20.0%Incorrect 68.0%Abstained 12.0%

Artificial Analysis benchmarks

GDPval-AA v2
16%
τ³-Banking
15%
Terminal-Bench v2.1
43%
SciCode
43%
Humanity’s Last Exam
24%
GPQA Diamond
86%
CritPt
1%
AA-Omniscience
26%
AA-LCR
68%

Similar models