readable AI benchmarks (simplified)
Alibaba
Qwen3.7 Max
Qwen3.7 Max has a Quality Score of 54.4, ranking 25th among 53 scored models. Per 1M tokens, pricing is $2.50 input, $7.50 output, $0.250 cached input, and $3.13 cache write. Its Reliability Score is 57.0, ranking 20th among 53 scored models, above average (total average: 50.1). Its Value Score is 36.1, around average compared with the total Value average of 39.6.
Quality Score54.4coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Score36.1
Reliability57.0
Cache Discount90%
Model specification
- Reasoning
- reasoning
- Input modalities
- none
- Output modalities
- none
- Context window
- not available
- Weights
- closed weights
Token prices USD per 1M tokens
- Input
- $2.50
- Output
- $7.50
- Cached Input
- $0.250
- Cache write
- $3.13
- Cost per task
- $1.03
Capability
- Intelligence
- 46.0
- Coding
- 66.0
- Agentic
- 30.6
- Omniscience
- 14.1
- Correct
- 30.1%
- Blended price
- $1.43
Answer outcomes
Attempt rate: 46.2%
Correct 30.1%Incorrect 16.0%Abstained 53.8%
Artificial Analysis benchmarks
- GDPval-AA v2
- 39%
- τ³-Banking
- 11%
- Terminal-Bench v2.1
- 75%
- SciCode
- 49%
- Humanity’s Last Exam
- 38%
- GPQA Diamond
- 92%
- CritPt
- 13%
- AA-Omniscience
- 57%
- AA-LCR
- 69%
Similar models
- Qwen3.7 PlusAlibabaQuality Score 43.6Value Score 36.1
- GPT-5.6 Terra (medium)OpenAIQuality Score 52.8Value Score 37.1
- Muse SparkMetaQuality Score 54.3Value Score not available
- GPT-5.6 Terra (high)OpenAIQuality Score 55.7Value Score 39.2
- GPT-5.6 Luna (xhigh)OpenAIQuality Score 53.0Value Score 50.2
- Kimi K2.6KimiQuality Score 51.9Value Score 38.5