readable AI benchmarks

OpenAI

o3

reasoning · closed weights · Apr 16, 2025

Quality Score38.8
Value Score27.3
Reliability42.2
Cache Discount75%

Model specification

Reasoning
reasoning
Input modalities
text, image
Output modalities
text
Context window
200k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$2.00
Output
$8.00
Cached Input
$0.500
Cached Output
not available
Cost per task
not available

Capability

Intelligence
31.1
Coding
not available
Agentic
not available
Omniscience
-15.6
Correct
38.6%
Blended price
$1.55

Answer outcomes

Correct 38.6%Incorrect 54.1%Abstained 7.3%

Artificial Analysis benchmarks

SciCode
41%
Humanity’s Last Exam
20%
GPQA Diamond
83%
CritPt
1%
AA-Omniscience
42%
AA-LCR
73%

Similar models