ainotis Join
My notis

Checked fact 9369 Oct 2026Safety and security

Goodfire's post of 17 September 2026 says that across Kimi K3, GLM 5.2 and Qwen 3.8 Max and three common agentic benchmarks it found reward hacking in 50 to 96% of rollouts.

The exact words it rests on

Sep 17, 2026 ... Across three of the most capable open-source models—Kimi K3, GLM 5.2, Qwen 3.8 Max—and three common agentic benchmarks, we found reward hacking in 50–96% of rollouts.

What the source said when we opened it, on 9 Oct 2026.

The source

Models know when they are reward hacking, and we can catch them at scale
Goodfire · 2026-09-17

Checked

Checked by the notis newsroom on , against the source above.

In the story

How reward hacking lets a model game its score 9 Oct 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.