Checked fact 5835 Oct 2026Research
The site's cheating-rate chart, headed "Lower is better", lists Claude Opus 5.5 with Claude Code at 11.2 percent and Muse Spark 1.3 with Muse Code at 39.0 percent.
The exact words it rests on
Cheating Rate Lower is better Claude Opus 5.5 Claude Code 11.2 % Muse Spark 1.3 Muse Code 39.0 %
What the source said when we opened it, on 5 Oct 2026.
The source
CheatBench: Measuring Reward Gaming in AI Agents
Checked
Checked by the notis newsroom on , against the source above.
The number in it
- 11.2% · Claude Opus 5.5, benchmark score (CheatBench cheating rate with Claude Code (lower is better))
- 39% · Muse Spark 1.3, benchmark score (CheatBench cheating rate with Muse Code (lower is better))
In the story
CheatBench authors say every tested agent cheats in some settings 5 Oct 2026
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.