readable AI benchmarks (simplified)

readable AI benchmarks

Google

Gemini 3.8 Flash (medium)

reasoning · closed weights · Sep 2, 2026

Gemini 3.8 Flash (medium) has a Quality Score of 70.6, ranking 17th among 275 scored models. Per 1M tokens, pricing is $0.750 input, $3.75 output, $0.075 cached input, and $0.750 cached output. It accepts text, image, video, and speech input, outputs text, and has a 1M token context window. Its Reliability Score is 64.3, ranking 15th among 275 scored models, above average (total average: 34.4). Its Value Score is 53.6, high compared with the total Value average of 28.6.

Quality Score70.6coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Score53.6
Reliability64.3
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
text, image, video, speech
Output modalities
text
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$0.750
Output
$3.75
Cached Input
$0.075
Cached Output
$0.750
Cost per task
$0.410

Capability

Intelligence
56.6
Coding
74.1
Agentic
48.5
Omniscience
28.6
Correct
53.0%
Blended price
$0.578

Answer outcomes

Attempt rate: 77.4%

Correct 53.0%Incorrect 24.4%Abstained 22.6%

Artificial Analysis benchmarks

GDPval-AA v2
50%
τ³-Banking
46%
Terminal-Bench v2.1
84%
SciCode
54%
Humanity’s Last Exam
42%
GPQA Diamond
94%
CritPt
12%
AA-Omniscience
64%
AA-LCR
82%

Similar models