readable AI benchmarks (simplified)
Google
Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning)
Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) has a Quality Score of 24.5, ranking 316th among 551 scored models. It has a 1M token context window. Its Reliability Score is 30.1, ranking 305th among 553 scored models, below average (total average: 35.6).
Quality Score24.5coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Scorenot available
Factual reliability30.1
Cache Discountnot available
Model specification
- Reasoning
- non-reasoning
- Input modalities
- none
- Output modalities
- none
- Context window
- 1M tokens
- Weights
- closed weights
Token prices USD per 1M tokens
- Input
- not available
- Output
- not available
- Cached Input
- not available
- Cache write
- not available
- Cost per task
- not available
Capability
- Intelligence
- 12.4
- Coding
- not available
- Agentic
- not available
- Omniscience
- -39.9
- Correct
- 26.8%
- Blended price
- not available
Answer outcomes
Fully graded outcomes (Correct + Incorrect): 93.6%
Correct 26.8%Incorrect 66.7%Partial / not attempted 6.4%
Artificial Analysis benchmarks
- Humanity’s Last Exam
- 9%
- GPQA Diamond
- 77%
- CritPt
- 0%
- AA-Omniscience
- 30%
- AA-LCR
- 60%
Similar models
- Gemma 4 31B (Non-reasoning)GoogleQuality Score 22.8Value Score 19.6
- Gemini 2.5 Flash (Non-reasoning)GoogleQuality Score 21.9Value Score 17.5
- Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning)GoogleQuality Score 21.5Value Score 19.9
- Gemini 2.0 Flash (Feb '25)GoogleQuality Score 21.2Value Score not available
- Gemma 4 26B A4B (Non-reasoning)GoogleQuality Score 19.5Value Score 16.8
- Gemma 4 E4B (Non-reasoning)GoogleQuality Score 20.5Value Score not available