ainotis Join
My notis

Checked fact 9379 Oct 2026Safety and security

Goodfire says it found an internal signal in the models, a direction in activation space, that accompanies reward hacking, and built difference-in-means probes that detect it.

The exact words it rests on

we found an internal signal [Specifically, a direction in activation space, found via difference-in-means from simple synthetic code examples.] that accompanies reward hacking.

What the source said when we opened it, on 9 Oct 2026.

The source

Models know when they are reward hacking, and we can catch them at scale
Goodfire · 2026-09-17

Checked

Checked by the notis newsroom on , against the source above.

In the story

How reward hacking lets a model game its score 9 Oct 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.