Skip to content
Benchmark

Epoch AI evaluates Muse Spark 1.3 (xhigh)

FrontierMath-Tiers-1-3-v2-Private 74.4% · FrontierMath-Tier-4-v2-Private 41.5% · Chess Puzzles 35.0% · Mystery Game Puzzles 14.1%

Type
BENCHMARK_RESULT
Detection
benchmark
Confidence
high
Status

Primary source

Epoch AI Benchmarking Hub · epoch.ai/benchmarks ↗

Independent benchmark runs (GPQA Diamond, SWE-bench Verified, FrontierMath, ...) plus externally reported leaderboards and the Epoch Capabilities Index (ECI).

Related events