readable AI benchmarks (simplified)

readable AI benchmarks

SpaceXAI

Grok 4

reasoning · closed weights · Jul 10, 2025

Grok 4 has a Quality Score of 42.8, ranking 145th among 551 scored models. Per 1M tokens, pricing is $3.00 input and $15.00 output. It has a 256k token context window. Its Reliability Score is 51.0, ranking 99th among 553 scored models, above average (total average: 35.6).

Quality Score42.8coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Scorenot available
Factual reliability51.0
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
256k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$3.00
Output
$15.00
Cached Input
not available
Cache write
not available
Cost per task
not available

Capability

Intelligence
22.5
Coding
not available
Agentic
not available
Omniscience
2.1
Correct
40.5%
Blended price
not available

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 78.9%

Correct 40.5%Incorrect 38.4%Partial / not attempted 21.1%

Artificial Analysis benchmarks

Humanity’s Last Exam
27%
GPQA Diamond
88%
CritPt
2%
AA-Omniscience
51%
AA-LCR
68%

Similar models

Support me! Patreon