readable AI benchmarks (simplified)
Meta
Muse Spark
Muse Spark has a Quality Score of 54.3, ranking 26th among 53 scored models. Its Reliability Score is 52.0, ranking 22nd among 53 scored models, above average (total average: 50.1).
Quality Score54.3coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Scorenot available
Reliability52.0
Cache Discountnot available
Model specification
- Reasoning
- reasoning
- Input modalities
- none
- Output modalities
- none
- Context window
- not available
- Weights
- closed weights
Token prices USD per 1M tokens
- Input
- not available
- Output
- not available
- Cached Input
- not available
- Cache write
- not available
- Cost per task
- not available
Capability
- Intelligence
- 43.1
- Coding
- 58.6
- Agentic
- 28.7
- Omniscience
- 4.1
- Correct
- 44.6%
- Blended price
- not available
Answer outcomes
Attempt rate: 85.2%
Correct 44.6%Incorrect 40.5%Abstained 14.8%
Artificial Analysis benchmarks
- GDPval-AA v2
- 32%
- τ³-Banking
- 20%
- Terminal-Bench v2.1
- 62%
- SciCode
- 52%
- Humanity’s Last Exam
- 40%
- GPQA Diamond
- 88%
- CritPt
- 11%
- AA-Omniscience
- 52%
- AA-LCR
- 70%
Similar models
- Muse Spark 1.1 (xhigh)MetaQuality Score 61.2Value Score 44.5
- Qwen3.7 MaxAlibabaQuality Score 54.4Value Score 36.1
- GPT-5.6 Terra (medium)OpenAIQuality Score 52.8Value Score 37.1
- Kimi K2.6KimiQuality Score 51.9Value Score 38.5
- DeepSeek V4 Pro (max)DeepSeekQuality Score 51.8Value Score 46.9
- GPT-5.5 (low)OpenAIQuality Score 57.6Value Score 34.6