ainotis Join
My notis

Checked fact 9389 Oct 2026Safety and security

Goodfire says that on Kimi K3 a combined probe and language-model monitor reduces the cost of language-model monitoring by 90% with only about a 1% drop in precision.

The exact words it rests on

On Kimi K3, a probe + LLM combined setup reduces the cost of LLM monitoring by 90% with only a ~1% drop in precision.

What the source said when we opened it, on 9 Oct 2026.

The source

Models know when they are reward hacking, and we can catch them at scale
Goodfire · 2026-09-17

Checked

Checked by the notis newsroom on , against the source above.

In the story

How reward hacking lets a model game its score 9 Oct 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.