readable AI benchmarks (simplified)
DeepSeek
DeepSeek V4 Flash 0420 (Reasoning, Max Effort)
DeepSeek V4 Flash 0420 (Reasoning, Max Effort) has a Quality Score of 37.6, ranking 194th among 553 scored models. Per 1M tokens, pricing is $0.130 input, $0.280 output, and $0.028 cached input. It has a 1M token context window. Its Reliability Score is 38.0, ranking 238th among 555 scored models, above average (total average: 35.7). Its Value Score is 34.5, around average compared with the total Value average of 30.7.
Quality Score37.6coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score34.5
Factual reliability38.0
Cache Discount78%
Model specification
- Reasoning
- reasoning
- Input modalities
- none
- Output modalities
- none
- Context window
- 1M tokens
- Weights
- open weights
Token prices USD per 1M tokens
- Input
- $0.130
- Output
- $0.280
- Cached Input
- $0.028
- Cache write
- not available
- Cost per task
- $0.118
Capability
- Intelligence
- 24.2
- Coding
- not available
- Agentic
- not available
- Omniscience
- -23.9
- Correct
- 36.8%
- Blended price
- $0.074
Answer outcomes
Fully graded outcomes (Correct + Incorrect): 97.5%
Correct 36.8%Incorrect 60.7%Partial / not attempted 2.5%
Artificial Analysis benchmarks
- GDPval-AA v2
- 26%
- τ³-Banking
- 31%
- SciCode
- 45%
- Humanity’s Last Exam
- 35%
- GPQA Diamond
- 89%
- CritPt
- 7%
- AA-Omniscience
- 38%
- AA-LCR
- 74%
Similar models
- DeepSeek V4 Flash 0420 (Reasoning, High Effort)DeepSeekQuality Score 39.2Value Score 36.0
- DeepSeek V3.2 (Reasoning)DeepSeekQuality Score 35.9Value Score 32.0
- DeepSeek V4 Pro 0424 (Reasoning, High Effort)DeepSeekQuality Score 45.5Value Score 38.8
- DeepSeek V3.2 SpecialeDeepSeekQuality Score 31.7Value Score not available
- DeepSeek V4 Pro 0424 (Reasoning, Max Effort)DeepSeekQuality Score 45.7Value Score 39.0
- DeepSeek V3.2 Exp (Reasoning)DeepSeekQuality Score 30.0Value Score 26.6