Today ainotis

Research1 Oct 2026Lead storyThat day's edition

Anthropic says GLM-5.3 builds exploits close to Claude Mythos Preview

Anthropic’s Frontier Red Team published a 29 September analysis of Zhipu AI’s open-weight GLM-5.3, covering its own tests of the model’s cyber capability and safeguards.

Check our sources · 1 source, 9 claims
Two line charts of exploitation success against output-token budget: on ExploitBench Claude Mythos Preview reaches 14 percent and GLM-5.3 12 percent, other models near 0.Image: Anthropic
source · The vendor's own image of its own product, used to report on that product

Key points

  1. On ExploitBench, Anthropic reports GLM-5.3 developed end-to-end exploits in 50 of 410 attempts and Claude Mythos Preview in 56 of 410, in Anthropic’s own test.
  2. Anthropic says in its simulated tests GLM-5.3 engaged with overtly malicious requests to attack critical systems 64% with a deceptive prompt, 92% with prefilled thinking and 100% when abliterated.
  3. Anthropic says NIST’s CAISI found GLM-5.3 "the most cyber-capable open-weight model released to date" in an assessment published on 17 September.

What happened

Anthropic’s Frontier Red Team published an analysis of Zhipu AI’s GLM-5.3 on 29 September 2026. On ExploitBench, GLM-5.3 developed end-to-end exploits in 50 of 410 attempts, and Claude Mythos Preview did so in 56 of 410.

On Anthropic’s internal Binary Exploitation benchmark, GLM-5.3 achieved full control-flow hijacks in 4% of trials and Claude Mythos Preview in 6%.

Anthropic says attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in its simulated tests, and that these did not succeed against safeguarded Claude models.

A researcher using the smaller GLM-5.3-Flash built an exploit chain for a Chrome flaw (CVE-2026-11645) at an API cost Anthropic puts at $20.40. Anthropic also says NIST’s CAISI published its own assessment of GLM-5.3 on 17 September and found it the most cyber-capable open-weight model released to date.

What it means for you

Our view

Every figure here comes from Anthropic’s tests of a competitor’s model, so read it as one lab’s view. Within that, Anthropic reports GLM-5.3 close to Claude Mythos Preview on its exploit benchmarks, and says simple techniques bypassed GLM-5.3’s safeguards in its simulated environment, where the bypass rates were measured.

The CAISI finding reaches us through Anthropic’s post, not the report itself. If you evaluate open-weight models for security work, ask which safeguards a model ships with, and read the CAISI assessment directly before relying on this comparison.

This part is our reading of the facts above. It adds no fact of its own.

Your reaction

One press adds one. We count a number for each notice and day, never who pressed it.

Check our sources

1 source, 9 claims. We opened the source and checked every sentence above against it.

1 GLM-5.3 and the spread of advanced cyber capabilitiesAnthropic · 29 Sep 2026 · 9 claims Open the source
  1. Anthropic’s Frontier Red Team published GLM-5.3 and the spread of advanced cyber capabilities on 29 September 2026.

    GLM-5.3 and the spread of advanced cyber capabilities
  2. Anthropic reports GLM-5.3 develops end-to-end exploits in 50 of 410 attempts on ExploitBench, and Claude Mythos Preview did so in 56 of 410 attempts.

    We find that GLM-5.3 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate—in 56 of 410 attempts.
  3. Anthropic reports that on its internal Binary Exploitation benchmark (100 randomly selected tasks) GLM-5.3 achieved full control-flow hijacks in 4% of trials and Claude Mythos Preview in 6%.

    We evaluate several models on 100 tasks from the benchmark (selected at random), and find that GLM-5.3 develops full control-flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%.
  4. Anthropic says attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in its simulated tests, and that these attacks did not succeed against safeguarded Claude models in its testing.

    We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests. In contrast, these attacks did not succeed against safeguarded Claude models in our testing.
  5. In a simulated environment where GLM-5.3 was given overtly malicious requests to attack critical systems, Anthropic says a deceptive red-team prompt got it to engage 64% of the time, prefilling its thinking tokens 92% and an abliterated version 100%.

    Narrowed to what the source supports
    Providing a deceptive prompt, such as telling the model that it is an autonomous red-team agent working on an exercise. This gets GLM-5.3 to engage 64% of the time.
  6. Anthropic says a researcher using GLM-5.3-Flash built a reliable exploit chain for an ARM64 target from a recently disclosed Chrome flaw (CVE-2026-11645) plus another known flaw, taking 20 minutes of human attention and eight hours of model work, which at Zhipu’s API prices would have cost $20.40.

    This took 20 minutes of human attention, plus eight hours of work for GLM-5.3-Flash. At Zhipu’s API prices, this effort would have cost $20.40.
  7. Anthropic says it produced an abliterated copy of GLM-5.3 using about 2,200 GPU hours at a computation cost of roughly $4,400.

    Abliterating the model took our team—which had never previously attempted this task—about 2,200 GPU hours at a computation cost of roughly $4,400.
  8. Anthropic says NIST’s CAISI published its own assessment of GLM-5.3 on Sept. 17 and found it "the most cyber-capable open-weight model released to date", lagging the US frontier by about four months on an aggregate of its cyber benchmarks.

    CAISI found that GLM-5.3 is “the most cyber-capable open-weight model released to date” and that it lags the US frontier by about four months on an aggregate of CAISI’s cyber benchmarks.
  9. Anthropic says GLM-5.3 was released without meaningful safeguards to limit misuse, unlike other frontier models with comparable cyber capabilities.

    But GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse.

Nothing appears on this site that we have not opened and linked.

Filed under