readable AI benchmarks

OpenAI

gpt-oss-120b (low)

reasoning · open weights · Aug 5, 2025

Quality Score15.1
Value Score16.6
Reliability23.3
Cache Discount0%

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
131,072 tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.150
Output
$0.595
Cached Input
$0.150
Cached Output
not available
Cost per task
$0.020

Capability

Intelligence
14.9
Coding
21.2
Agentic
1.0
Omniscience
-53.5
Correct
19.8%
Blended price
$0.195

Answer outcomes

Correct 19.8%Incorrect 73.3%Abstained 6.9%

Artificial Analysis benchmarks

GDPval-AA v2
0%
τ³-Banking
3%
Terminal-Bench v2.1
14%
SciCode
36%
Humanity’s Last Exam
6%
GPQA Diamond
67%
CritPt
0%
AA-Omniscience
23%
AA-LCR
46%

Similar models