readable AI benchmarks

SpaceXAI

Grok 4.6 (low)

reasoning · closed weights · Aug 12, 2026

Quality Score65.1
Value Score43.2
Reliability63.0
Cache Discount75%

Model specification

Reasoning
reasoning
Input modalities
text, image
Output modalities
text
Context window
500k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$2.00
Output
$6.00
Cached Input
$0.500
Cached Output
not available
Cost per task
$0.221

Capability

Intelligence
51.7
Coding
66.3
Agentic
47.8
Omniscience
25.9
Correct
43.3%
Blended price
$1.35

Answer outcomes

Correct 43.3%Incorrect 17.4%Abstained 39.3%

Artificial Analysis benchmarks

GDPval-AA v2
53%
τ³-Banking
38%
Terminal-Bench v2.1
75%
SciCode
48%
Humanity’s Last Exam
28%
GPQA Diamond
88%
CritPt
6%
AA-Omniscience
63%
AA-LCR
79%

Similar models