readable AI benchmarks (simplified)

readable AI benchmarks

OpenAI

gpt-oss-20b (High)

reasoning · open weights · Aug 5, 2025

gpt-oss-20b (High) has a Quality Score of 16.1, ranking 480th among 571 scored models. Per 1M tokens, pricing is $0.070 input and $0.180 output. It accepts text input, outputs text, and has a 131,072 token context window. Its Reliability Score is 18.5, ranking 502nd among 573 scored models, below average (total average: 36.5).

Quality Score16.1coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Scorenot available
Factual reliability18.5
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
131,072 tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.070
Output
$0.180
Cached Input
not available
Cache write
not available
Cost per task
$0.012

Capability

Intelligence
9.0
Coding
not available
Agentic
not available
Omniscience
-63.0
Correct
16.0%
Blended price
not available

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 95.0%

Correct 16.0%Incorrect 79.0%Partial / not attempted 4.9%

Artificial Analysis benchmarks

GDPval-AA v2
0%
τ³-Banking
7%
SciCode
39%
Humanity’s Last Exam
11%
GPQA Diamond
69%
CritPt
1%
AA-Omniscience
18%
AA-LCR
35%

Similar models

Support me! Patreon