ainotis Join
My notis

Checked fact 9399 Oct 2026Safety and security

Goodfire says its probes catch 3.1% more hacks in Kimi K3 but 7.9% fewer in GLM 5.2 than chain-of-thought monitors on DeepSWE at a matched false positive rate.

The exact words it rests on

Compared to chain-of-thought monitors, our probes catch 3.1% more hacks in Kimi K3 but 7.9% fewer hacks in GLM 5.2 on DeepSWE at a matched false positive rate.

What the source said when we opened it, on 9 Oct 2026.

The source

Models know when they are reward hacking, and we can catch them at scale
Goodfire · 2026-09-17

Checked

Checked by the notis newsroom on , against the source above.

In the story

How reward hacking lets a model game its score 9 Oct 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.