readable AI benchmarks

Meta

Llama 4 Maverick

non-reasoning · open weights · Apr 5, 2025

Quality Score17.6
Value Score16.9
Reliability29.1
Cache Discount6%

Model specification

Reasoning
non-reasoning
Input modalities
text, image
Output modalities
text
Context window
1M tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.260
Output
$0.910
Cached Input
$0.245
Cached Output
not available
Cost per task
$0.035

Capability

Intelligence
14.5
Coding
16.3
Agentic
1.2
Omniscience
-41.9
Correct
24.9%
Blended price
$0.314

Answer outcomes

Correct 24.9%Incorrect 66.8%Abstained 8.4%

Artificial Analysis benchmarks

GDPval-AA v2
0%
τ³-Banking
4%
Terminal-Bench v2.1
8%
SciCode
33%
Humanity’s Last Exam
5%
GPQA Diamond
67%
CritPt
0%
AA-Omniscience
29%
AA-LCR
50%

Similar models