readable AI benchmarks (simplified)

readable AI benchmarks

DeepSeek

DeepSeek V3.1 (Non-reasoning)

non-reasoning · open weights · Aug 21, 2025

DeepSeek V3.1 (Non-reasoning) has a Quality Score of 24.9, ranking 309th among 551 scored models. Per 1M tokens, pricing is $0.570 input, $1.68 output, and $0.570 cached input. It accepts text input, outputs text, and has a 128k token context window. Its Reliability Score is 28.6, ranking 324th among 553 scored models, below average (total average: 35.6). Its Value Score is 18.2, low compared with the total Value average of 30.7.

Quality Score24.9coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score18.2
Factual reliability28.6
Cache Discount0%

Model specification

Reasoning
non-reasoning
Input modalities
text
Output modalities
text
Context window
128k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.570
Output
$1.68
Cached Input
$0.570
Cache write
not available
Cost per task
not available

Capability

Intelligence
13.7
Coding
not available
Agentic
not available
Omniscience
-42.8
Correct
23.1%
Blended price
$0.681

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 89.0%

Correct 23.1%Incorrect 65.9%Partial / not attempted 11.0%

Artificial Analysis benchmarks

Humanity’s Last Exam
7%
GPQA Diamond
74%
CritPt
0%
AA-Omniscience
29%
AA-LCR
47%

Similar models

Support me! Patreon