readable AI benchmarks (simplified)
OpenAI
GPT-5.4 mini (xhigh)
GPT-5.4 mini (xhigh) has a Quality Score of 44.6, ranking 38th among 53 scored models. Per 1M tokens, pricing is $0.750 input, $4.50 output, and $0.075 cached input. Its Reliability Score is 40.7, ranking 44th among 53 scored models, below average (total average: 50.1). Its Value Score is 36.1, around average compared with the total Value average of 39.6.
Quality Score44.6coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Score36.1
Reliability40.7
Cache Discount90%
Model specification
- Reasoning
- reasoning
- Input modalities
- none
- Output modalities
- none
- Context window
- not available
- Weights
- closed weights
Token prices USD per 1M tokens
- Input
- $0.750
- Output
- $4.50
- Cached Input
- $0.075
- Cache write
- not available
- Cost per task
- $0.452
Capability
- Intelligence
- 40.0
- Coding
- 56.1
- Agentic
- 30.2
- Omniscience
- -18.7
- Correct
- 37.5%
- Blended price
- $0.653
Answer outcomes
Attempt rate: 93.6%
Correct 37.5%Incorrect 56.1%Abstained 6.4%
Artificial Analysis benchmarks
- GDPval-AA v2
- 33%
- τ³-Banking
- 21%
- Terminal-Bench v2.1
- 59%
- SciCode
- 50%
- Humanity’s Last Exam
- 27%
- GPQA Diamond
- 87%
- CritPt
- 10%
- AA-Omniscience
- 41%
- AA-LCR
- 69%
Similar models
- GPT-5.6 Luna (medium)OpenAIQuality Score 43.9Value Score 41.0
- GPT-5.6 Terra (low)OpenAIQuality Score 48.4Value Score 33.7
- GPT-5.6 Luna (low)OpenAIQuality Score 40.0Value Score 36.9
- GPT-5.6 Luna (high)OpenAIQuality Score 50.5Value Score 47.6
- GPT-5.6 Terra (medium)OpenAIQuality Score 52.8Value Score 37.1
- GPT-5.6 Luna (xhigh)OpenAIQuality Score 53.0Value Score 50.2