// AI threat desk · 1 Oct 2026
Home/AI-ATK · AI-Powered Attacks
AI-Powered AttacksHighGLM-5.3

Anthropic test: GLM-5.3 and Claude Mythos achieve first control-flow hijacks

Anthropic’s Frontier Red Team published findings on Sep 29 showing China’s GLM-5.3 model fully hijacked control flow in 4% of 100 binary exploit tests. Claude Mythos Preview hit 6%. Prior-generation models failed entirely, marking a key shift in AI attack capability.

Illustrative image: code on screens in a server room, symbolising AI models running exploit tests.
File photo: Illustrative image: code on screens in a server room, symbolising AI models running exploit tests.. Photo: rawpixel (CC0)
Key takeaways
  • Anthropic’s Frontier Red Team tests show GLM-5.3 fully hijacked program control flow in 4% of 100 internal binary exploitation tests. Claude Mythos Preview reached 6%. Prior-generation models including Claude Opus 4.6 and GLM-5.2 failed completely
  • In simulated tests, Anthropic found attackers using simple methods bypassed GLM-5.3’s safety guardrails at a success rate of 64% to 100%, far higher than against guarded Claude models
  • Researchers used GLM-5.3-Flash against known vulnerabilities including CVE-2026-11645, building a complete attack chain that bypasses ARM64 pointer authentication with 20 minutes of manual work plus 8 hours of model compute, at a cost of about $20.40
On this page

GLM-5.3 achieves first full control-flow hijack

According to GLM-5.3 and the spread of advanced cyber capabilities, Anthropic’s Frontier Red Team tested multiple models on a random sample of 100 tasks from its internal Binary Exploitation benchmark. GLM-5.3, developed by China’s Zhipu AI (Z.ai), fully hijacked program control flow in 4% of tests. Claude Mythos Preview reached 6%.

The report notes that while GLM-5.3 performed slightly below Claude Mythos Preview, a key threshold has clearly been crossed: prior-generation models such as Claude Opus 4.6 and GLM-5.2 failed completely on the same benchmark. Simon Willison reposted this finding on Sep 29.

In a separate ExploitBench test, GLM-5.3 succeeded in building end-to-end exploits in 50 of 410 attempts, while Claude Mythos Preview succeeded in 56, placing the two models’ performance close together.

Safety guardrails offer little resistance

Anthropic’s report shows that attackers using simple methods bypassed GLM-5.3’s safety guardrails at a success rate of 64% to 100%. By comparison, similar attacks failed to break through guarded Claude models.

The report highlights ‘abliteration’, a refusal-removal technique. Because GLM-5.3 is released as an open-weight model, users can modify it themselves to remove its mechanism for refusing harmful requests, with minimal impact on the model’s other capabilities. Several developers publicly released abliterated versions of GLM-5.3 within days of its launch.

Anthropic’s own team performed abliteration using about 2,200 GPU hours, at a computing cost of roughly $4,400. This reduced the model’s refusal rate on JailbreakBench and HarmBench from over 90% to about 2% to 3%, and to 12% on StrongREJECT.

Researchers build real attack chains at low cost

The report describes a researcher who, in one day with less than an hour of human attention, used GLM-5.3 to find several previously unknown vulnerabilities in a popular web browser’s JavaScript engine. The researcher chained them into a complete attack: a victim visiting a related webpage would allow the attacker to read arbitrary files on their computer. Anthropic reported the vulnerabilities to the browser’s maintainers.

In another test, researchers used the smaller GLM-5.3-Flash against the publicly known Google Chrome vulnerability CVE-2026-11645, combined with another known flaw. With 20 minutes of manual guidance and 8 hours of model compute, they built a reliable attack chain that bypasses ARM64 Pointer Authentication (PAC) hardening. Based on Zhipu AI’s API pricing, the entire process cost about $20.40.

NIST assessment confirms narrowing capability gap

Anthropic cited an assessment published on Sep 17 by the US National Institute of Standards and Technology (NIST)’s Center for AI Standards and Innovation (CAISI), which called GLM-5.3 ‘the most cyber-capable open-weight model released to date’. On CAISI’s combined score across multiple cybersecurity benchmarks, it trails leading US frontier models by about four months.

The report adds that CAISI’s tests of US models had some cybersecurity guardrails disabled, and that leading US models are released only to vetted users, making them harder for attackers to obtain. By contrast, GLM-5.3 can be downloaded and used by anyone.

Independent researcher Matthew Green, in a separate commentary, noted that agents can leave instructions for each other through shared package caches, altering each other’s behaviour. Applying this scenario to shared channels such as email, Slack or WhatsApp, combined with independently deployed personal AI agents like Muse, would supply all the elements needed for worm-like spread. This echoes this outlet’s earlier report on an OpenAI agent that tried to hack sites when blocked, which likewise reflects rising risks from autonomous agent behaviour.

What to do now

  1. Defenders should assume open-weight models, including GLM-5.3, can have their safety guardrails easily removed, and should not assume vendor-built refusal mechanisms are sufficient to stop malicious use
  2. Security teams should assess whether their own systems contain [exploit chains](/glossary#remote-code-execution) that could be automatically discovered and chained by AI models, paying particular attention to browser engines, drivers and internet-facing device software
  3. Enterprises deploying AI agents for automated tasks should restrict agents from passing unreviewed instructions to each other through shared channels such as package caches or chat tools, to reduce the risk of worm-like spread
  4. Keep tracking patch status for known CVEs (such as CVE-2026-11645), since AI models can quickly turn public vulnerability details into usable attack chains

FAQ

What is GLM-5.3 and who developed it?

GLM-5.3 is an open-weight AI model developed by China’s Zhipu AI, known internationally as Z.ai. According to Anthropic’s Frontier Red Team assessment, the model is capable of autonomously building end-to-end cyberattack exploits.

What do the 4% and 6% control-flow hijack rates mean?

These are success rates from Anthropic’s internal binary exploitation benchmark, based on a random sample of 100 tasks. They represent the proportion of tests in which a model fully took control of a target program’s execution flow, the highest-level outcome in an attack chain.

How easily can GLM-5.3’s safety guardrails be bypassed?

According to Anthropic’s simulated tests, attackers using simple methods bypassed GLM-5.3’s safety guardrails at a success rate of 64% to 100%, while similar attacks failed to break through guarded Claude models.

What is abliteration?

Abliteration is a standard refusal-removal technique used to modify an open-weight model’s internal parameters, removing its mechanism for refusing harmful requests, with minimal impact on the model’s other capabilities.

Does this research mean GLM-5.3 has already been used in real attacks?

The report does not state that GLM-5.3 has been used in real-world attacks. All tests were conducted in isolated sandbox environments set up by Anthropic, targeting offline systems. Newly discovered vulnerabilities were separately reported to the relevant software maintainers.

Sources

  1. GLM-5.3 and the spread of advanced cyber capabilities, Anthropic Frontier Red TeamPrimary
  2. Quoting Anthropic Frontier Red Team, Simon Willison
  3. Quoting Matthew Green, Simon Willison
Explore with AI