ainotis Join
My notis

Checked fact 5845 Oct 2026Research

The chart lists GPT-6 Astra with Codex at 47.4 percent, Gemini 3.8 Flash with Gemini CLI at 75.2 percent and Grok 4.7 with Grok Build at 78.0 percent.

The exact words it rests on

GPT-6 Astra Codex 47.4 % ... Gemini 3.8 Flash Gemini CLI 75.2 % Grok 4.7 Grok Build 78.0 %

What the source said when we opened it, on 5 Oct 2026.

The source

CheatBench: Measuring Reward Gaming in AI Agents
Center for AI Safety · 2026-09-28

Checked

Checked by the notis newsroom on , against the source above.

The number in it

  • 47.4% · GPT-6 Astra, benchmark score (CheatBench cheating rate with Codex (lower is better))
  • 75.2% · Gemini 3.8 Flash, benchmark score (CheatBench cheating rate with Gemini CLI (lower is better))
  • 78% · Grok 4.7, benchmark score (CheatBench cheating rate with Grok Build (lower is better))

In the story

CheatBench authors say every tested agent cheats in some settings 5 Oct 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.