readable AI benchmarks (simplified)
Sapiens AI
Agnes 2.5 Pro Alpha (based on Qwen3.5 397B A17B)
Agnes 2.5 Pro Alpha (based on Qwen3.5 397B A17B) has a Quality Score of 39.3, ranking 193rd among 571 scored models. Per 1M tokens, pricing is $0.450 input, $0.900 output, and $0.010 cached input. It has a 1M token context window. Its Reliability Score is 37.5, ranking 259th among 573 scored models, above average (total average: 36.5). Its Value Score is 33.4, around average compared with the total Value average of 31.4.
Quality Score39.3coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score33.4
Factual reliability37.5
Cache Discount98%
Model specification
- Reasoning
- reasoning
- Input modalities
- none
- Output modalities
- none
- Context window
- 1M tokens
- Weights
- open weights
Token prices USD per 1M tokens
- Input
- $0.450
- Output
- $0.900
- Cached Input
- $0.010
- Cache write
- not available
- Cost per task
- not available
Capability
- Intelligence
- 26.8
- Coding
- not available
- Agentic
- not available
- Omniscience
- -25.1
- Correct
- 33.5%
- Blended price
- $0.187
Answer outcomes
Fully graded outcomes (Correct + Incorrect): 92.1%
Correct 33.5%Incorrect 58.6%Partial / not attempted 7.9%
Artificial Analysis benchmarks
- τ³-Banking
- 12%
- SciCode
- 43%
- Humanity’s Last Exam
- 34%
- GPQA Diamond
- 88%
- CritPt
- 11%
- AA-Omniscience
- 37%
- AA-LCR
- 76%
Similar models
- Agnes 2.5 Pro BetaSapiens AIQuality Score 49.5Value Score 46.1
- Agnes 3.0 FlashSapiens AIQuality Score 49.7Value Score 47.7
- Apodex 1.1ApodexQuality Score 39.8Value Score 31.3
- Qwen3.8 27B (Low)AlibabaQuality Score 38.5Value Score 29.5
- Hy3TencentQuality Score 39.8Value Score 35.6
- Nex-N2-Pro (Based on Qwen3.5-397B-A17B)Nex AGIQuality Score 39.9Value Score not available