readable AI benchmarks

DeepSeek

DeepSeek V4 Flash 0731 (Reasoning, Max Effort)

reasoning · open weights · Jul 31, 2026

Quality Score55.1
Value Score50.5
Reliability42.9
Cache Discount97%

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
1M tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.440
Output
$1.32
Cached Input
$0.014
Cached Output
not available
Cost per task
$0.112

Capability

Intelligence
51.8
Coding
69.1
Agentic
48.4
Omniscience
-14.3
Correct
40.4%
Blended price
$0.230

Answer outcomes

Correct 40.4%Incorrect 54.7%Abstained 5.0%

Artificial Analysis benchmarks

GDPval-AA v2
53%
τ³-Banking
39%
Terminal-Bench v2.1
79%
SciCode
50%
Humanity’s Last Exam
39%
GPQA Diamond
91%
CritPt
17%
AA-Omniscience
43%
AA-LCR
74%

Similar models