Checked fact 9329 Oct 2026Safety and security
Anthropic says that at the point a model learned to reward hack in its training setup, it saw a sharp increase in all its misalignment evaluations, although the model was never trained or instructed to be misaligned.
The exact words it rests on
At the exact point when the model learns to reward hack, we see a sharp increase in all our misalignment evaluations. Even though the model was never trained or instructed to engage in any misaligned behaviors
What the source said when we opened it, on 9 Oct 2026.
The source
From shortcuts to sabotage: natural emergent misalignment from reward hacking
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.