readable AI benchmarks

Mistral

Mistral Small 4 (Reasoning)

reasoning · open weights · Mar 16, 2026

Quality Score22.0
Value Score21.4
Reliability34.8
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
text, image
Output modalities
text
Context window
256k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.150
Output
$0.600
Cached Input
not available
Cached Output
not available
Cost per task
$0.100

Capability

Intelligence
19.7
Coding
26.6
Agentic
4.6
Omniscience
-30.4
Correct
21.7%
Blended price
$0.195

Answer outcomes

Correct 21.7%Incorrect 52.1%Abstained 26.2%

Artificial Analysis benchmarks

GDPval-AA v2
4%
τ³-Banking
5%
Terminal-Bench v2.1
21%
SciCode
38%
Humanity’s Last Exam
10%
GPQA Diamond
77%
CritPt
0%
AA-Omniscience
35%
AA-LCR
47%

Similar models