readable AI benchmarks (simplified)

readable AI benchmarks

OpenAI

GPT-5.4 nano (xhigh)

reasoning · closed weights · Mar 17, 2026

GPT-5.4 nano (xhigh) has a Quality Score of 33.6, ranking 223rd among 551 scored models. Per 1M tokens, pricing is $0.200 input, $1.25 output, and $0.020 cached input. It accepts text and image input, outputs text, and has a 400k token context window. Its Reliability Score is 35.3, ranking 255th among 553 scored models, below average (total average: 35.6). Its Value Score is 28.6, around average compared with the total Value average of 30.7.

Quality Score33.6coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score28.6
Factual reliability35.3
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
text, image
Output modalities
text
Context window
400k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$0.200
Output
$1.25
Cached Input
$0.020
Cache write
not available
Cost per task
$0.183

Capability

Intelligence
20.7
Coding
not available
Agentic
not available
Omniscience
-29.5
Correct
25.7%
Blended price
$0.179

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 80.8%

Correct 25.7%Incorrect 55.1%Partial / not attempted 19.2%

Artificial Analysis benchmarks

GDPval-AA v2
22%
τ³-Banking
27%
SciCode
47%
Humanity’s Last Exam
28%
GPQA Diamond
82%
CritPt
9%
AA-Omniscience
35%
AA-LCR
77%

Similar models

Support me! Patreon