readable AI benchmarks (simplified)

readable AI benchmarks

Mistral

Mistral Medium 3.5

reasoning · open weights · Apr 29, 2026

Mistral Medium 3.5 has a Quality Score of 26.7, ranking 312th among 577 scored models. Per 1M tokens, pricing is $1.50 input, $7.50 output, and $0.150 cached input. It accepts text and image input, outputs text, and has a 256k token context window. Its Reliability Score is 31.6, ranking 316th among 579 scored models, below average (total average: 36.6). Its Value Score is 18.4, low compared with the total Value average of 31.8.

Quality Score26.7coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score18.4
Factual reliability31.6
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
text, image
Output modalities
text
Context window
256k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$1.50
Output
$7.50
Cached Input
$0.150
Cache write
not available
Cost per task
$0.502

Capability

Intelligence
14.2
Coding
not available
Agentic
not available
Omniscience
-36.8
Correct
24.7%
Blended price
$1.16

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 86.1%

Correct 24.7%Incorrect 61.5%Partial / not attempted 13.9%

Artificial Analysis benchmarks

GDPval-AA v2
13%
τ³-Banking
15%
SciCode
40%
Humanity’s Last Exam
14%
GPQA Diamond
75%
CritPt
0%
AA-Omniscience
32%
AA-LCR
69%

Similar models

Support me! Patreon