The figures here are Anthropic’s and NIST CAISI’s own. Anthropic sells competing models and ran its tests itself. This publication verified none of it independently and did not obtain a response from Z.ai.
Anthropic says attackers can bypass the safeguards on Z.ai’s open-weight GLM-5.3 between 64% and 100% of the time in its simulated tests. Anthropic sells competing models, ran the tests itself and published the results on 29 September 2026. This publication read that post and NIST’s separate assessment, verified neither independently, and did not obtain a response from Z.ai.
What Anthropic Published
Anthropic’s research post of 29 September 2026 analyses GLM-5.3, the model from Z.ai, formerly Zhipu AI. It says GLM-5.3, like Claude Mythos Preview, “has strong capabilities for autonomously building end-to-end cyber exploits,” and that it has been released without meaningful safeguards to limit misuse.
Anthropic says that in its simulated tests attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques, and that these attacks did not succeed against safeguarded Claude models in its testing.
Anthropic sells competing models, the tests are its own and ran in sandboxed or simulated environments. This publication did not obtain a response from Z.ai, so none is included.

What NIST’s CAISI Found
NIST’s Center for AI Standards and Innovation says Z.ai released GLM-5.3 on 14 August 2026 and publicly released its weights two weeks later. CAISI’s key findings were that GLM-5.3 is “the most cyber-capable open-weight model released to date” and that its cyber capabilities are “significantly lower than those of current U.S. frontier models.”
CAISI says GLM-5.3 lags the U.S. frontier by about four months on an aggregate of its cyber benchmarks. Its table gives GLM-5.3 at 40.4% (74 of 183) on SEC-Bench Pro against a U.S. frontier best of 90.2%, and 61.1% on ExploitBench against 100.0%.
On ExploitGym (userspace) CAISI gives GLM-5.3 9.4% against 44.4%, and on OSS-Fuzz 7.7% against 23.2%. CAISI says its U.S. frontier includes trusted-access releases and that U.S. models were tested with cyber safeguards disabled where applicable.
What Anthropic Measured on Capability
On ExploitBench, Anthropic says GLM-5.3 developed end-to-end exploits in 50 of 410 attempts, and Claude Mythos Preview in 56 of 410. In its internal Binary Exploitation benchmark, on 100 tasks chosen at random, it says GLM-5.3 reached full control-flow hijacks in 4% of trials and Mythos Preview in 6%.
Anthropic adds that earlier models, Claude Opus 4.6 and GLM-5.2, did not succeed on any of those tasks. This publication’s own arithmetic: 50 of 410 is 12.2% and 56 of 410 is 13.7%.
Anthropic says its capability findings “broadly match” CAISI’s. The two comparisons use different yardsticks, CAISI the best U.S. score on each benchmark and Anthropic one named model, so the figures are not directly comparable.
The Safeguard Tests
Anthropic says GLM-5.3 refused all overtly harmful requests in its simulated world when asked directly. It then lists three ways it got the model to engage. A deceptive prompt posing as a red-team exercise got engagement 64% of the time, prefilling the model’s thinking tokens 92%, and an abliterated copy 100%.
Anthropic says each cell had 50 samples and that none of the techniques got safeguarded Claude models to carry out the tasks. It says the Anthropic API gives no way to prefill Claude’s thinking and that Claude’s weights, not being public, cannot be abliterated.
Abliteration is a refusal-reduction technique. Anthropic says its team, which had not tried it before, produced an abliterated GLM-5.3 in about 2,200 GPU hours at a computation cost of roughly $4,400. It says the edit moved refusal rates from above 90% to about 3% on JailbreakBench, about 2% on HarmBench and 12% on StrongREJECT, with general science scores on GPQA-Diamond unchanged.
This publication’s own arithmetic: $4,400 over 2,200 GPU hours is about $2 per GPU hour.
What Anthropic Says It Means
Anthropic says GLM-5.3 gives malicious actors access to capabilities without meaningful restrictions, which it calls “a meaningful step change.” It says it thinks it likely that state and non-state actors will use models like it to cause real-world harm, and that governments should conduct safety testing on sufficiently capable models.
Anthropic also says defenders benefit from the same capabilities, and that it is working to expand access to Claude’s cyber capabilities to as many defenders as it can. Those are Anthropic’s views, from a company that sells such access.
What Is Not Established
Neither Anthropic’s nor CAISI’s tests have been repeated by this publication, and Z.ai’s own account of GLM-5.3’s safeguards is not included. Whether the 64% to 100% rates carry over outside Anthropic’s simulated world is not stated in the post.
Nothing here says what any actor has done with the model. All figures are Anthropic’s or CAISI’s own and none was independently verified. Nothing in this piece is investment advice.
This piece draws on Anthropic’s research post of 29 September 2026 and NIST CAISI’s assessment of GLM-5.3. Anthropic sells competing models and ran its own tests, and CAISI’s figures are its own. This publication verified none of it independently and did not obtain a response from Z.ai. The 12.2% and 13.7% figures and the $2 per GPU hour figure are this publication’s own arithmetic. Nothing above predicts anything, and nothing here is investment advice.
Sources: anthropic.com · nist.gov · anthropic.com · NIST CAISI — ‘CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities’, September 2026









