readable AI benchmarks (simplified)

readable AI benchmarks

SpaceXAI

Grok 4.3 (Non-reasoning)

non-reasoning · closed weights · release date not listed

Grok 4.3 (Non-reasoning) has a Quality Score of 27.5, ranking 299th among 571 scored models. Per 1M tokens, pricing is $1.25 input, $2.50 output, $0.200 cached input, and $1.25 cache write. It has a 1M token context window. Its Reliability Score is 33.5, ranking 291st among 573 scored models, below average (total average: 36.5). Its Value Score is 20.3, low compared with the total Value average of 31.4.

Quality Score27.5coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score20.3
Factual reliability33.5
Cache Discount84%

Model specification

Reasoning
non-reasoning
Input modalities
none
Output modalities
none
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$1.25
Output
$2.50
Cached Input
$0.200
Cache write
$1.25
Cost per task
$0.170

Capability

Intelligence
14.0
Coding
not available
Agentic
not available
Omniscience
-33.1
Correct
23.7%
Blended price
$0.640

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 80.5%

Correct 23.7%Incorrect 56.8%Partial / not attempted 19.5%

Artificial Analysis benchmarks

GDPval-AA v2
22%
τ³-Banking
8%
SciCode
39%
Humanity’s Last Exam
7%
GPQA Diamond
66%
CritPt
0%
AA-Omniscience
33%
AA-LCR
32%

Similar models

Support me! Patreon