readable AI benchmarks (simplified)

readable AI benchmarks

SpaceXAI

Grok 4.3 (high)

reasoning · closed weights · Apr 30, 2026

Grok 4.3 (high) has a Quality Score of 48.6, ranking 95th among 551 scored models. Per 1M tokens, pricing is $1.25 input, $2.50 output, $0.200 cached input, and $1.25 cache write. It accepts text, image, and pdf input, outputs text, and has a 1M token context window. Its Reliability Score is 59.0, ranking 59th among 553 scored models, above average (total average: 35.6). Its Value Score is 35.9, high compared with the total Value average of 30.7.

Quality Score48.6coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score35.9
Factual reliability59.0
Cache Discount84%

Model specification

Reasoning
reasoning
Input modalities
text, image, pdf
Output modalities
text
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$1.25
Output
$2.50
Cached Input
$0.200
Cache write
$1.25
Cost per task
$0.165

Capability

Intelligence
24.9
Coding
not available
Agentic
not available
Omniscience
18.0
Correct
34.8%
Blended price
$0.640

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 51.6%

Correct 34.8%Incorrect 16.8%Partial / not attempted 48.4%

Artificial Analysis benchmarks

GDPval-AA v2
21%
τ³-Banking
12%
SciCode
48%
Humanity’s Last Exam
37%
GPQA Diamond
90%
CritPt
8%
AA-Omniscience
59%
AA-LCR
73%

Similar models

Support me! Patreon