readable AI benchmarks (simplified)
Sapiens AI
Agnes 2.5 Pro Beta
Agnes 2.5 Pro Beta has a Quality Score of 44.8, ranking 60th among 263 scored models. Per 1M tokens, pricing is $0.100 input, $0.300 output, and $0.010 cached input. It accepts text and image input, outputs text, and has a 1M token context window. Its Reliability Score is 44.7, ranking 71st among 263 scored models, above average (total average: 33.3). Its Value Score is 47.9, high compared with the total Value average of 27.5.
Quality Score44.8coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Score47.9
Reliability44.7
Cache Discount90%
Model specification
- Reasoning
- reasoning
- Input modalities
- text, image
- Output modalities
- text
- Context window
- 1M tokens
- Weights
- closed weights
Token prices USD per 1M tokens
- Input
- $0.100
- Output
- $0.300
- Cached Input
- $0.010
- Cached Output
- not available
- Cost per task
- not available
Capability
- Intelligence
- 49.1
- Coding
- 62.3
- Agentic
- 43.8
- Omniscience
- -10.5
- Correct
- 16.8%
- Blended price
- $0.057
Answer outcomes
Attempt rate: 44.1%
Correct 16.8%Incorrect 27.3%Abstained 55.9%
Artificial Analysis benchmarks
- GDPval-AA v2
- 48%
- τ³-Banking
- 36%
- Terminal-Bench v2.1
- 70%
- SciCode
- 48%
- Humanity’s Last Exam
- 38%
- GPQA Diamond
- 91%
- CritPt
- 16%
- AA-Omniscience
- 45%
- AA-LCR
- 78%
Similar models
- Agnes 2.5 Pro AlphaSapiens AIQuality Score 41.1Value Score 40.2
- MiniMax-M3MiniMaxQuality Score 46.6Value Score 40.2
- Qwen3.8 27B (xhigh)AlibabaQuality Score 47.2Value Score 41.5
- Hy3TencentQuality Score 44.0Value Score 44.2
- Solar Pro 4UpstageQuality Score 43.9Value Score 37.9
- Inkling SmallThinking MachinesQuality Score 46.0Value Score 41.1