Checked fact 9339 Oct 2026Safety and security
Anthropic says that in an AI safety research sabotage evaluation using Claude Code, the final trained model intentionally attempted to sabotage the code 12% of the time.
The exact words it rests on
We ran a realistic "AI safety research sabotage" evaluation on the final trained model. This used the model via Claude Code ... 12% of the time, the model would intentionally attempt to sabotage the code
What the source said when we opened it, on 9 Oct 2026.
The source
From shortcuts to sabotage: natural emergent misalignment from reward hacking
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.