readable AI benchmarks (simplified)

readable AI benchmarks

SpaceXAI

Grok 4.3 (Medium)

reasoning · closed weights · release date not listed

Grok 4.3 (Medium) has a Quality Score of 48.2, ranking 118th among 571 scored models. Per 1M tokens, pricing is $1.25 input, $2.50 output, $0.200 cached input, and $1.25 cache write. It has a 1M token context window. Its Reliability Score is 58.4, ranking 74th among 573 scored models, above average (total average: 36.5). Its Value Score is 35.6, around average compared with the total Value average of 31.4.

Quality Score48.2coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score35.6
Factual reliability58.4
Cache Discount84%

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$1.25
Output
$2.50
Cached Input
$0.200
Cache write
$1.25
Cost per task
not available

Capability

Intelligence
24.8
Coding
not available
Agentic
not available
Omniscience
16.7
Correct
28.8%
Blended price
$0.640

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 40.8%

Correct 28.8%Incorrect 12.1%Partial / not attempted 59.2%

Artificial Analysis benchmarks

Humanity’s Last Exam
30%
GPQA Diamond
89%
CritPt
5%
AA-Omniscience
58%
AA-LCR
75%

Similar models

Support me! Patreon