readable AI benchmarks (simplified)

readable AI benchmarks

Z AI

GLM-5.1 (Reasoning)

reasoning · open weights · Apr 7, 2026

GLM-5.1 (Reasoning) has a Quality Score of 45.3, ranking 131st among 551 scored models. Per 1M tokens, pricing is $1.29 input, $4.07 output, and $0.260 cached input. It accepts text input, outputs text, and has a 200k token context window. Its Reliability Score is 50.4, ranking 106th among 553 scored models, above average (total average: 35.6). Its Value Score is 32.3, around average compared with the total Value average of 30.7.

Quality Score45.3coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score32.3
Factual reliability50.4
Cache Discount80%

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
200k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$1.29
Output
$4.07
Cached Input
$0.260
Cache write
not available
Cost per task
$1.02

Capability

Intelligence
26.1
Coding
not available
Agentic
not available
Omniscience
0.8
Correct
23.7%
Blended price
$0.846

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 46.5%

Correct 23.7%Incorrect 22.8%Partial / not attempted 53.4%

Artificial Analysis benchmarks

GDPval-AA v2
30%
τ³-Banking
14%
SciCode
45%
Humanity’s Last Exam
30%
GPQA Diamond
87%
CritPt
5%
AA-Omniscience
50%
AA-LCR
74%

Similar models

Support me! Patreon