readable AI benchmarks

SpaceXAI

Grok 4.3 (Non-reasoning)

non-reasoning · closed weights · Apr 30, 2026

Quality Score27.8
Value Score22.7
Reliability33.5
Cache Discount84%

Model specification

Reasoning
non-reasoning
Input modalities
text, image
Output modalities
text
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$1.25
Output
$2.50
Cached Input
$0.200
Cached Output
$1.25
Cost per task
$0.295

Capability

Intelligence
25.0
Coding
35.2
Agentic
23.0
Omniscience
-33.1
Correct
23.7%
Blended price
$0.640

Answer outcomes

Correct 23.7%Incorrect 56.8%Abstained 19.5%

Artificial Analysis benchmarks

GDPval-AA v2
30%
τ³-Banking
8%
Terminal-Bench v2.1
34%
SciCode
37%
Humanity’s Last Exam
7%
GPQA Diamond
66%
CritPt
0%
AA-Omniscience
33%
AA-LCR
28%

Similar models