readable AI benchmarks (simplified)

readable AI benchmarks

Meta

Muse Spark

reasoning · closed weights · Apr 8, 2026

Muse Spark has a Quality Score of 54.3, ranking 26th among 53 scored models. Its Reliability Score is 52.0, ranking 22nd among 53 scored models, above average (total average: 50.1).

Quality Score54.3coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Scorenot available
Reliability52.0
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
none
Output modalities
none
Context window
not available
Weights
closed weights

Token prices USD per 1M tokens

Input
not available
Output
not available
Cached Input
not available
Cache write
not available
Cost per task
not available

Capability

Intelligence
43.1
Coding
58.6
Agentic
28.7
Omniscience
4.1
Correct
44.6%
Blended price
not available

Answer outcomes

Attempt rate: 85.2%

Correct 44.6%Incorrect 40.5%Abstained 14.8%

Artificial Analysis benchmarks

GDPval-AA v2
32%
τ³-Banking
20%
Terminal-Bench v2.1
62%
SciCode
52%
Humanity’s Last Exam
40%
GPQA Diamond
88%
CritPt
11%
AA-Omniscience
52%
AA-LCR
70%

Similar models