ainotis Join
My notis

Checked fact 3001 Oct 2026Research

Anthropic reports that on its internal Binary Exploitation benchmark (100 randomly selected tasks) GLM-5.3 achieved full control-flow hijacks in 4% of trials and Claude Mythos Preview in 6%.

The exact words it rests on

We evaluate several models on 100 tasks from the benchmark (selected at random), and find that GLM-5.3 develops full control-flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%.

What the source said when we opened it, on 1 Oct 2026.

The source

GLM-5.3 and the spread of advanced cyber capabilities
Anthropic · 2026-09-29

Checked

Checked by the notis newsroom on , against the source above.

The number in it

  • 4% · glm-5-3, benchmark score (Binary Exploitation (Anthropic internal, 100 tasks))
  • 6% · claude-mythos-preview, benchmark score (Binary Exploitation (Anthropic internal, 100 tasks))

In the story

Anthropic says GLM-5.3 builds exploits close to Claude Mythos Preview 1 Oct 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.