Checked fact 3001 Oct 2026Research
Anthropic reports that on its internal Binary Exploitation benchmark (100 randomly selected tasks) GLM-5.3 achieved full control-flow hijacks in 4% of trials and Claude Mythos Preview in 6%.
The exact words it rests on
We evaluate several models on 100 tasks from the benchmark (selected at random), and find that GLM-5.3 develops full control-flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%.
What the source said when we opened it, on 1 Oct 2026.
The source
GLM-5.3 and the spread of advanced cyber capabilities
Checked
Checked by the notis newsroom on , against the source above.
The number in it
- 4% · glm-5-3, benchmark score (Binary Exploitation (Anthropic internal, 100 tasks))
- 6% · claude-mythos-preview, benchmark score (Binary Exploitation (Anthropic internal, 100 tasks))
In the story
Anthropic says GLM-5.3 builds exploits close to Claude Mythos Preview 1 Oct 2026
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.