Checked fact 9359 Oct 2026Safety and security
Anthropic recommends inoculation prompting, which recasts reward hacking as an acceptable behaviour in the training prompt, as a practical mitigation, and says it has already started using the technique.
The exact words it rests on
We recommend inoculation prompting using language such as that as a practical mitigation that AI developers could adopt ... and we have already started making use of this technique in training Claude.
What the source said when we opened it, on 9 Oct 2026.
The source
From shortcuts to sabotage: natural emergent misalignment from reward hacking
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.