readable AI benchmarks (simplified)
Anthropic
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) has a Quality Score of 75.7, ranking 1st among 282 scored models. Per 1M tokens, pricing is $4.00 input, $20.00 output, $0.200 cached input, and $5.00 cache write. It accepts text, image, and pdf input, outputs text, and has a 1M token context window. Its Reliability Score is 73.2, ranking 1st among 282 scored models, above average (total average: 36.9). Its Value Score is 47.9, high compared with the total Value average of 30.3.
Quality Score75.7coverage 78.8%: missing Epoch General ECI, Coding, CursorBench 3.2, Agentic and DeepSWE v1.1
Value Score47.9
Reliability73.2
Cache Discount95%
Model specification
- Reasoning
- reasoning
- Input modalities
- text, image, pdf
- Output modalities
- text
- Context window
- 1M tokens
- Weights
- closed weights
Token prices USD per 1M tokens
- Input
- $4.00
- Output
- $20.00
- Cached Input
- $0.200
- Cache write
- $5.00
- Cost per task
- $5.98
Capability
- Intelligence
- 57.6
- Coding
- not available
- Agentic
- not available
- Omniscience
- 46.4
- Correct
- 66.2%
- Blended price
- $2.94
Answer outcomes
Attempt rate: 86.0%
Correct 66.2%Incorrect 19.8%Abstained 14.0%
Artificial Analysis benchmarks
- GDPval-AA v2
- 67%
- SciCode
- 67%
- Humanity’s Last Exam
- 61%
- CritPt
- 32%
- AA-Omniscience
- 73%
- AA-LCR
- 85%
Similar models
- Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)AnthropicQuality Score 73.8Value Score 46.6
- Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)AnthropicQuality Score 74.1Value Score 41.2
- Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)AnthropicQuality Score 73.5Value Score 40.9
- Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)AnthropicQuality Score 72.1Value Score 45.0
- Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)AnthropicQuality Score 72.0Value Score 39.6
- Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)AnthropicQuality Score 71.0Value Score 43.6