readable AI benchmarks (simplified)
OpenAI
o1
o1 has a Quality Score of 34.0, ranking 218th among 551 scored models. Per 1M tokens, pricing is $15.00 input, $60.00 output, and $7.50 cached input. It accepts text, image, and pdf input, outputs text, and has a 200k token context window. Its Reliability Score is 44.5, ranking 181st among 553 scored models, above average (total average: 35.6). Its Value Score is 17.6, low compared with the total Value average of 30.7.
Quality Score34.0coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score17.6
Factual reliability44.5
Cache Discount50%
Model specification
- Reasoning
- reasoning
- Input modalities
- text, image, pdf
- Output modalities
- text
- Context window
- 200k tokens
- Weights
- closed weights
Token prices USD per 1M tokens
- Input
- $15.00
- Output
- $60.00
- Cached Input
- $7.50
- Cache write
- not available
- Cost per task
- not available
Capability
- Intelligence
- 15.2
- Coding
- not available
- Agentic
- not available
- Omniscience
- -11.1
- Correct
- 34.5%
- Blended price
- $14.25
Answer outcomes
Fully graded outcomes (Correct + Incorrect): 80.1%
Correct 34.5%Incorrect 45.6%Partial / not attempted 19.9%
Artificial Analysis benchmarks
- Humanity’s Last Exam
- 7%
- GPQA Diamond
- 75%
- CritPt
- 0%
- AA-Omniscience
- 44%
- AA-LCR
- 65%
Similar models
- GPT-5 mini (high)OpenAIQuality Score 33.6Value Score 27.4
- GPT-5.4 nano (xhigh)OpenAIQuality Score 33.6Value Score 28.6
- GPT-5.4 mini (medium)OpenAIQuality Score 35.0Value Score 25.8
- GPT-5.4 nano (medium)OpenAIQuality Score 35.9Value Score 30.6
- o3OpenAIQuality Score 36.7Value Score 24.3
- GPT-5.1 Codex mini (high)OpenAIQuality Score 36.6Value Score 29.9