readable AI benchmarks (simplified)

readable AI benchmarks

Meta

Muse Spark 1.3 (max)

reasoning · closed weights · Sep 2, 2026

Muse Spark 1.3 (max) has a Quality Score of 71.5, ranking 15th among 277 scored models. It accepts text, image, and video input, outputs text, and has a 1M token context window. Its Reliability Score is 62.5, ranking 22nd among 277 scored models, above average (total average: 34.6).

Quality Score71.5coverage 91.2%: missing Epoch General ECI, CursorBench 3.2 and DeepSWE v1.1
Value Scorenot available
Reliability62.5
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
text, image, video
Output modalities
text
Context window
1M tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
not available
Output
not available
Cached Input
not available
Cached Output
not available
Cost per task
not available

Capability

Intelligence
62.1
Coding
76.3
Agentic
59.3
Omniscience
24.9
Correct
43.8%
Blended price
not available

Answer outcomes

Attempt rate: 62.7%

Correct 43.8%Incorrect 18.9%Abstained 37.3%

Artificial Analysis benchmarks

GDPval-AA v2
63%
τ³-Banking
52%
Terminal-Bench v2.1
86%
SciCode
57%
Humanity’s Last Exam
49%
GPQA Diamond
94%
CritPt
25%
AA-Omniscience
62%
AA-LCR
79%

Similar models