readable AI benchmarks (simplified)

readable AI benchmarks

Mistral

Mistral Small 4 (Reasoning)

reasoning · open weights · Mar 16, 2026

Mistral Small 4 (Reasoning) has a Quality Score of 26.1, ranking 314th among 571 scored models. Per 1M tokens, pricing is $0.150 input, $0.600 output, and $0.015 cached input. It accepts text and image input, outputs text, and has a 256k token context window. Its Reliability Score is 34.8, ranking 281st among 573 scored models, below average (total average: 36.5). Its Value Score is 23.4, low compared with the total Value average of 31.4.

Quality Score26.1coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score23.4
Factual reliability34.8
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
text, image
Output modalities
text
Context window
256k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.150
Output
$0.600
Cached Input
$0.015
Cache write
not available
Cost per task
$0.015

Capability

Intelligence
11.3
Coding
not available
Agentic
not available
Omniscience
-30.4
Correct
21.7%
Blended price
$0.100

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 73.8%

Correct 21.7%Incorrect 52.1%Partial / not attempted 26.2%

Artificial Analysis benchmarks

GDPval-AA v2
0%
τ³-Banking
5%
SciCode
39%
Humanity’s Last Exam
10%
GPQA Diamond
77%
CritPt
0%
AA-Omniscience
35%
AA-LCR
50%

Similar models

Support me! Patreon