readable AI benchmarks (simplified)

readable AI benchmarks

Google

Gemini 3.5 Flash (high)

reasoning · closed weights · May 19, 2026

Gemini 3.5 Flash (high) has a Quality Score of 55.4, ranking 69th among 551 scored models. Per 1M tokens, pricing is $1.50 input, $9.00 output, and $0.150 cached input. It accepts text, image, video, audio, and pdf input, outputs text, and has a 1M token context window. Its Reliability Score is 60.6, ranking 49th among 553 scored models, above average (total average: 35.6). Its Value Score is 37.5, high compared with the total Value average of 30.7.

Quality Score55.4coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score37.5
Factual reliability60.6
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
text, image, video, audio, pdf
Output modalities
text
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$1.50
Output
$9.00
Cached Input
$0.150
Cache write
not available
Cost per task
$1.56

Capability

Intelligence
32.6
Coding
not available
Agentic
not available
Omniscience
21.2
Correct
51.4%
Blended price
$1.31

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 81.6%

Correct 51.4%Incorrect 30.2%Partial / not attempted 18.4%

Artificial Analysis benchmarks

GDPval-AA v2
34%
τ³-Banking
32%
SciCode
54%
Humanity’s Last Exam
43%
GPQA Diamond
92%
CritPt
13%
AA-Omniscience
61%
AA-LCR
73%

Similar models

Support me! Patreon