readable AI benchmarks (simplified)

readable AI benchmarks

SpaceXAI

Grok 3

non-reasoning · closed weights · Feb 19, 2025

Grok 3 has a Quality Score of 26.0, ranking 297th among 551 scored models. Per 1M tokens, pricing is $4.00 input, $20.00 output, and $1.54 cached input. It has a 1M token context window. Its Reliability Score is 33.4, ranking 273rd among 553 scored models, below average (total average: 35.6). Its Value Score is 15.5, low compared with the total Value average of 30.7.

Quality Score26.0coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score15.5
Factual reliability33.4
Cache Discount62%

Model specification

Reasoning
non-reasoning
Input modalities
none
Output modalities
none
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$4.00
Output
$20.00
Cached Input
$1.54
Cache write
not available
Cost per task
not available

Capability

Intelligence
12.1
Coding
not available
Agentic
not available
Omniscience
-33.2
Correct
28.9%
Blended price
$3.88

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 91.0%

Correct 28.9%Incorrect 62.1%Partial / not attempted 9.0%

Artificial Analysis benchmarks

Humanity’s Last Exam
4%
GPQA Diamond
69%
CritPt
0%
AA-Omniscience
33%
AA-LCR
58%

Similar models

Support me! Patreon