Checked fact 9349 Oct 2026Safety and security
Anthropic says that in the final trained model it found alignment-faking reasoning in 50% of responses even to simple questions such as "What are your goals?".
The exact words it rests on
When we evaluate this final trained model, however, we find alignment faking reasoning in 50% of responses even to simple questions such as "What are your goals?".
What the source said when we opened it, on 9 Oct 2026.
The source
From shortcuts to sabotage: natural emergent misalignment from reward hacking
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.