readable AI benchmarks (simplified)

readable AI benchmarks

Mistral

Mistral Large 4 Preview

reasoning · closed weights · Oct 6, 2026

Mistral Large 4 Preview has a Quality Score of 53.2, ranking 90th among 572 scored models. Per 1M tokens, pricing is $1.36 input, $4.18 output, $0.140 cached input, and $1.36 cache write. It has a 524,288 token context window. Its Reliability Score is 47.4, ranking 158th among 574 scored models, above average (total average: 36.5). Its Value Score is 38.3, high compared with the total Value average of 31.5.

Quality Score53.2coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score38.3
Factual reliability47.4
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
524,288 tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$1.36
Output
$4.18
Cached Input
$0.140
Cache write
$1.36
Cost per task
$1.13

Capability

Intelligence
38.4
Coding
not available
Agentic
not available
Omniscience
-5.3
Correct
25.8%
Blended price
$0.788

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 56.9%

Correct 25.8%Incorrect 31.1%Partial / not attempted 43.1%

Artificial Analysis benchmarks

GDPval-AA v2
46%
SciCode
54%
Humanity’s Last Exam
35%
CritPt
11%
AA-Omniscience
47%
AA-LCR
81%

Similar models

Support me! Patreon