readable AI benchmarks

OpenAI

gpt-oss-120b (high)

reasoning · open weights · Aug 5, 2025

Quality Score22.4
Value Score24.3
Reliability25.4
Cache Discount0%

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
131,072 tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.150
Output
$0.600
Cached Input
$0.150
Cached Output
not available
Cost per task
$0.073

Capability

Intelligence
24.1
Coding
30.4
Agentic
13.4
Omniscience
-49.3
Correct
21.8%
Blended price
$0.195

Answer outcomes

Correct 21.8%Incorrect 71.0%Abstained 7.2%

Artificial Analysis benchmarks

GDPval-AA v2
15%
τ³-Banking
13%
Terminal-Bench v2.1
26%
SciCode
39%
Humanity’s Last Exam
20%
GPQA Diamond
78%
CritPt
1%
AA-Omniscience
25%
AA-LCR
51%

Similar models