Z.ai’s publicly released GLM-5.3 has been assessed as the most cyber-capable open-weight model tested to date, materially lowering barriers to vulnerability research and exploit development. Anthropic reported that the model autonomously produced end-to-end V8 exploits in 50 of 410 attempts and achieved control-flow hijacking in 4% of sampled OSS-Fuzz tasks; researcher-guided tests also found previously unknown JavaScript-engine flaws in a popular browser’s Linux build and chained them to read arbitrary files. In a separate test, GLM-5.3-Flash developed an ARM64 Chrome exploit chain involving CVE-2026-11645, bypassing pointer-authentication hardening after roughly eight hours of model work and 20 minutes of human involvement.
The model’s public weights, released two weeks after its August launch, enable broad redistribution and modification. Anthropic found its safety controls could be bypassed through deceptive prompting, thinking-token prefilling, and weight-level “abliteration,” with simulated harmful-task engagement rising to 64%, 92%, and 100%, respectively; abliterated variants were publicly released soon after launch. NIST’s CAISI assessed GLM-5.3 as ahead of prior open-weight leader Kimi K3 but estimated it remains about four months behind current U.S. frontier models, underscoring that its risk stems from the accessibility of high-end offensive capability rather than parity with the leading closed models.

Track how attackers are adapting to this technology.
9 events from the most recent confirmed update back to the earliest known activity.
NIST’s Center for AI Standards and Innovation stated that GLM-5.3 was the most cyber-capable open-weight model released to date.
Z.ai, formerly Zhipu AI, released the GLM-5.3 open-weight AI model.
Developers publicly released refusal-removed, or abliterated, versions of GLM-5.3 within days of the model’s release.
Z.ai publicly released GLM-5.3’s downloadable model weights approximately two weeks after the model launch.
Anthropic found GLM-5.3’s harmful-request safeguards could be bypassed through deceptive prompts or thinking-token prefilling, and removed through weight-level abliteration. In simulated malicious-task tests, engagement reached 64% with a cover story, 92% with prefilling, and 100% for an abliterated model.
Using public information supplied by a researcher, GLM-5.3-Flash developed an ARM64 exploit chain combining Chrome CVE-2026-11645 with another known flaw and bypassing pointer-authentication hardening. The exercise used about eight hours of model processing and approximately 20 minutes of researcher attention.
In a sandboxed researcher-guided session, GLM-5.3 identified several previously unknown flaws in a popular browser’s JavaScript engine on Linux and chained them into a malicious webpage capable of reading arbitrary files, including a demonstrated SSH private key. Anthropic disclosed the flaws to the relevant maintainer.
In Anthropic’s ExploitBench testing, GLM-5.3 generated end-to-end exploits for known V8 vulnerabilities in 50 of 410 attempts. It also achieved full control-flow hijacking in 4% of 100 sampled OSS-Fuzz binary-exploitation tasks.
CAISI evaluated GLM-5.3 across four vulnerability-discovery and exploit-development benchmarks, finding it the most capable open-weight model it had evaluated. CAISI nevertheless estimated it lagged the U.S. frontier by roughly four months in aggregate cyber capability.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.