readable AI benchmarks (simplified)

readable AI benchmarks

Mistral

Mistral Medium 3

non-reasoning · closed weights · May 7, 2025

Mistral Medium 3 has a Quality Score of 24.1, ranking 318th among 551 scored models. Per 1M tokens, pricing is $0.400 input and $2.00 output. It accepts text and image input, outputs text, and has a 128k token context window. Its Reliability Score is 34.3, ranking 266th among 553 scored models, below average (total average: 35.6).

Quality Score24.1coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Scorenot available
Factual reliability34.3
Cache Discountnot available

Model specification

Reasoning
non-reasoning
Input modalities
text, image
Output modalities
text
Context window
128k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$0.400
Output
$2.00
Cached Input
not available
Cache write
not available
Cost per task
not available

Capability

Intelligence
9.0
Coding
not available
Agentic
not available
Omniscience
-31.4
Correct
18.3%
Blended price
not available

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 68.1%

Correct 18.3%Incorrect 49.7%Partial / not attempted 31.9%

Artificial Analysis benchmarks

Humanity’s Last Exam
4%
GPQA Diamond
58%
CritPt
0%
AA-Omniscience
34%
AA-LCR
31%

Similar models

Support me! Patreon