readable AI benchmarks (simplified)

readable AI benchmarks

SpaceXAI

Grok 4.7 (high)

reasoning · closed weights · Sep 21, 2026

Grok 4.7 (high) has a Quality Score of 62.0, ranking 22nd among 264 scored models. It has a 500k token context window. Its Reliability Score is 65.5, ranking 17th among 264 scored models, above average (total average: 35.4).

Quality Score62.0coverage 78.8%: missing Epoch General ECI, Coding, CursorBench 3.2, Agentic and DeepSWE v1.1
Value Scorenot available
Reliability65.5
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
500k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
not available
Output
not available
Cached Input
not available
Cache write
not available
Cost per task
not available

Capability

Intelligence
46.3
Coding
not available
Agentic
not available
Omniscience
30.9
Correct
47.8%
Blended price
not available

Answer outcomes

Attempt rate: 64.7%

Correct 47.8%Incorrect 16.9%Abstained 35.3%

Artificial Analysis benchmarks

GDPval-AA v2
60%
SciCode
58%
Humanity’s Last Exam
42%
CritPt
18%
AA-Omniscience
65%
AA-LCR
77%

Similar models

Support me! Ko-Fi