readable AI benchmarks

Google

Gemma 4 31B (Non-reasoning)

non-reasoning · open weights · Apr 2, 2026

Quality Score19.8
Value Score21.9
Reliability24.2
Cache Discount7%

Model specification

Reasoning
non-reasoning
Input modalities
text, image, video
Output modalities
text
Context window
256k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.150
Output
$0.400
Cached Input
$0.140
Cached Output
not available
Cost per task
$0.044

Capability

Intelligence
22.3
Coding
33.2
Agentic
11.1
Omniscience
-51.7
Correct
16.6%
Blended price
$0.168

Answer outcomes

Correct 16.6%Incorrect 68.3%Abstained 15.1%

Artificial Analysis benchmarks

GDPval-AA v2
12%
τ³-Banking
9%
Terminal-Bench v2.1
29%
SciCode
41%
Humanity’s Last Exam
12%
GPQA Diamond
76%
CritPt
0%
AA-Omniscience
24%
AA-LCR
43%

Similar models