ainotis Join
My notis

Checked fact 9359 Oct 2026Safety and security

Anthropic recommends inoculation prompting, which recasts reward hacking as an acceptable behaviour in the training prompt, as a practical mitigation, and says it has already started using the technique.

The exact words it rests on

We recommend inoculation prompting using language such as that as a practical mitigation that AI developers could adopt ... and we have already started making use of this technique in training Claude.

What the source said when we opened it, on 9 Oct 2026.

The source

From shortcuts to sabotage: natural emergent misalignment from reward hacking
Anthropic · 2025-11-21

Checked

Checked by the notis newsroom on , against the source above.

In the story

How reward hacking lets a model game its score 9 Oct 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.