readable AI benchmarks (simplified)

readable AI benchmarks

OpenAI

GPT-5.4 mini (xhigh)

reasoning · closed weights · Mar 17, 2026

GPT-5.4 mini (xhigh) has a Quality Score of 44.6, ranking 38th among 53 scored models. Per 1M tokens, pricing is $0.750 input, $4.50 output, and $0.075 cached input. Its Reliability Score is 40.7, ranking 44th among 53 scored models, below average (total average: 50.1). Its Value Score is 36.1, around average compared with the total Value average of 39.6.

Quality Score44.6coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Score36.1
Reliability40.7
Cache Discount90%

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
not available
Weights
closed weights

Token prices USD per 1M tokens

Input
$0.750
Output
$4.50
Cached Input
$0.075
Cache write
not available
Cost per task
$0.452

Capability

Intelligence
40.0
Coding
56.1
Agentic
30.2
Omniscience
-18.7
Correct
37.5%
Blended price
$0.653

Answer outcomes

Attempt rate: 93.6%

Correct 37.5%Incorrect 56.1%Abstained 6.4%

Artificial Analysis benchmarks

GDPval-AA v2
33%
τ³-Banking
21%
Terminal-Bench v2.1
59%
SciCode
50%
Humanity’s Last Exam
27%
GPQA Diamond
87%
CritPt
10%
AA-Omniscience
41%
AA-LCR
69%

Similar models