readable AI benchmarks (simplified)

readable AI benchmarks

Tencent

Hy3

reasoning · open weights · Jul 6, 2026

Hy3 has a Quality Score of 39.8, ranking 189th among 571 scored models. Per 1M tokens, pricing is $0.136 input, $0.555 output, and $0.034 cached input. It accepts text input, outputs text, and has a 256k token context window. Its Reliability Score is 40.8, ranking 232nd among 573 scored models, above average (total average: 36.5). Its Value Score is 35.6, around average compared with the total Value average of 31.4.

Quality Score39.8coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score35.6
Factual reliability40.8
Cache Discount75%

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
256k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.136
Output
$0.555
Cached Input
$0.034
Cache write
not available
Cost per task
$0.072

Capability

Intelligence
25.3
Coding
not available
Agentic
not available
Omniscience
-18.5
Correct
32.0%
Blended price
$0.107

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 82.4%

Correct 32.0%Incorrect 50.4%Partial / not attempted 17.6%

Artificial Analysis benchmarks

GDPval-AA v2
28%
τ³-Banking
23%
SciCode
49%
Humanity’s Last Exam
33%
GPQA Diamond
90%
CritPt
5%
AA-Omniscience
41%
AA-LCR
79%

Similar models

Support me! Patreon