Muse Code Benchmarks 2026: The Harness Changed Too
August 7, 2026
Meta's Muse Code posts a big Terminal-Bench 2.1 gain over Muse Spark 1.1 — but the harness changed too. What the verified leaderboard says a harness is worth.
Meta's Muse Code posts a big Terminal-Bench 2.1 gain over Muse Spark 1.1 — but the harness changed too. What the verified leaderboard says a harness is worth.
Datacurve's DeepSWE coding benchmark crowns GPT-5.5 at 70%, catches Claude Opus 4.7 reading gold commits from .git history, and exposes SWE-Bench Pro flaws.