readable AI benchmarks (simplified)

readable AI benchmarks

Alibaba

Qwen3 Coder Next

non-reasoning · open weights · Feb 3, 2026

Qwen3 Coder Next has a Quality Score of 16.5, ranking 478th among 577 scored models. Per 1M tokens, pricing is $0.350 input, $1.20 output, and $0.350 cached input. It accepts text input, outputs text, and has a 256k token context window. Its Reliability Score is 18.8, ranking 500th among 579 scored models, below average (total average: 36.6). Its Value Score is 12.8, low compared with the total Value average of 31.8.

Quality Score16.5coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score12.8
Factual reliability18.8
Cache Discount0%

Model specification

Reasoning
non-reasoning
Input modalities
text
Output modalities
text
Context window
256k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.350
Output
$1.20
Cached Input
$0.350
Cache write
not available
Cost per task
$0.552

Capability

Intelligence
9.2
Coding
not available
Agentic
not available
Omniscience
-62.4
Correct
16.2%
Blended price
$0.435

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 94.7%

Correct 16.2%Incorrect 78.5%Partial / not attempted 5.3%

Artificial Analysis benchmarks

GDPval-AA v2
1%
τ³-Banking
5%
SciCode
36%
Humanity’s Last Exam
10%
GPQA Diamond
74%
CritPt
0%
AA-Omniscience
19%
AA-LCR
47%

Similar models

Support me! Patreon