Checked fact 9379 Oct 2026Safety and security
Goodfire says it found an internal signal in the models, a direction in activation space, that accompanies reward hacking, and built difference-in-means probes that detect it.
The exact words it rests on
we found an internal signal [Specifically, a direction in activation space, found via difference-in-means from simple synthetic code examples.] that accompanies reward hacking.
What the source said when we opened it, on 9 Oct 2026.
The source
Models know when they are reward hacking, and we can catch them at scale
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.