Checked fact 9309 Oct 2026Safety and security
Anthropic's post of 21 November 2025 defines reward hacking as an AI fooling its training process into assigning a high reward without actually completing the intended task.
The exact words it rests on
Nov 21, 2025 ... "reward hacking": an AI fooling its training process into assigning a high reward, without actually completing the intended task
What the source said when we opened it, on 9 Oct 2026.
The source
From shortcuts to sabotage: natural emergent misalignment from reward hacking
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.