readable AI benchmarks (simplified)

readable AI benchmarks

SpaceXAI

Grok 4.7 (xhigh)

reasoning · closed weights · Sep 21, 2026

Grok 4.7 (xhigh) has a Quality Score of 62.3, ranking 19th among 264 scored models. It accepts text, image, and pdf input, outputs text, and has a 500k token context window. Its Reliability Score is 66.0, ranking 14th among 264 scored models, above average (total average: 35.4).

Quality Score62.3coverage 78.8%: missing Epoch General ECI, Coding, CursorBench 3.2, Agentic and DeepSWE v1.1
Value Scorenot available
Reliability66.0
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
text, image, pdf
Output modalities
text
Context window
500k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
not available
Output
not available
Cached Input
not available
Cache write
not available
Cost per task
not available

Capability

Intelligence
46.4
Coding
not available
Agentic
not available
Omniscience
32.0
Correct
47.4%
Blended price
not available

Answer outcomes

Attempt rate: 62.9%

Correct 47.4%Incorrect 15.4%Abstained 37.1%

Artificial Analysis benchmarks

GDPval-AA v2
60%
SciCode
57%
Humanity’s Last Exam
43%
CritPt
18%
AA-Omniscience
66%
AA-LCR
77%

Similar models

Support me! Ko-Fi