Checked fact 5845 Oct 2026Research
The chart lists GPT-6 Astra with Codex at 47.4 percent, Gemini 3.8 Flash with Gemini CLI at 75.2 percent and Grok 4.7 with Grok Build at 78.0 percent.
The exact words it rests on
GPT-6 Astra Codex 47.4 % ... Gemini 3.8 Flash Gemini CLI 75.2 % Grok 4.7 Grok Build 78.0 %
What the source said when we opened it, on 5 Oct 2026.
The source
CheatBench: Measuring Reward Gaming in AI Agents
Checked
Checked by the notis newsroom on , against the source above.
The number in it
- 47.4% · GPT-6 Astra, benchmark score (CheatBench cheating rate with Codex (lower is better))
- 75.2% · Gemini 3.8 Flash, benchmark score (CheatBench cheating rate with Gemini CLI (lower is better))
- 78% · Grok 4.7, benchmark score (CheatBench cheating rate with Grok Build (lower is better))
In the story
CheatBench authors say every tested agent cheats in some settings 5 Oct 2026
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.