Beijing-based Z.ai released GLM-5.3 on Friday while holding back the model's downloadable weights and gating its most sensitive cybersecurity functions. In launch tests, GLM-5.3 scored 84.5% on CyberGym and edged Anthropic's restricted Mythos 5. Z.ai, which published downloadable weights for its previous GLM models, is gating this one with the kind of control American labs have used on their strongest cyber models.
GLM-5.3 uses the same base model as GLM-5.2. Z.ai said every reported gain came from a month of expanded post-training, with more task environments, a broader mix of work and more computing time. "As we scaled post-training, cyber capability developed faster than we expected," the company said. It had deliberately added vulnerability-discovery work, but said the model progressed from finding isolated flaws toward planning complete exploitation chains.
What Changed
- GLM-5.3 scored 84.5% on CyberGym, ahead of the 83.8% Z.ai reported for Anthropic's Mythos 5 and 83.6% for GPT-5.6 Sol.
- The lead does not survive the move from finding flaws to exploiting them: 54.4% on ExploitBench against 78.0% for Mythos 5.
- Z.ai is holding the downloadable weights until around August 28, the first time it has delayed a GLM weight release.
- Its vulnerability ledger lists 2,436 findings across 269 open-source projects, with 53 disclosed and 2,383 still under embargo.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
The 84.5% CyberGym result, up from GLM-5.2's 77.2%, covered 1,507 tasks from 188 software projects in Z.ai's August 14 evaluation. The company put Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6% on the same test. The advantage disappeared as the work moved toward exploitation. GLM-5.3 scored 54.4% on ExploitBench, up from GLM-5.2's 24.4%, but far below the 78.0% Z.ai reported for Mythos 5. GLM-5.3 completed 105 ExploitGym tasks under a two-hour normalized budget and 130 under six hours. Mythos 5 completed 181 and 247.
Anthropic released Mythos 5 in June only to verified private partners. "As this capability carries the greatest potential for misuse in security, we are limiting initial access to a small number of partners through Project Glasswing," Anthropic said.
Z.ai's vulnerability ledger supplies evidence outside benchmark tasks, though the company controls that record too. As of August 14, it listed 2,436 findings across 269 open-source projects after expert review, screening and deduplication. The total included 107 critical flaws and 990 rated high. Only 53 had been disclosed, while 2,383 remained under embargo. The affected software included the Linux kernel, Redis, WebKit and FreeBSD. The oldest flaw was introduced in 1981, and the listed vulnerabilities had gone undiscovered for an average of 26.6 years.
The cyber scores and the in-house Code Bench results are Z.ai's own, produced in Z.ai's own configuration, and no independent evaluator has replicated the cyber results. Artificial Analysis had not added the model as of August 14. On CyberGym, the model ran at maximum reasoning effort, got a single attempt per task and had no time limit on any task. ExploitGym's time budgets were rescaled with model throughput rates rather than measured wall-clock time. Z.ai's private Code Bench cannot be audited outside the company.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
Z.ai shares fell on August 14, the day of the launch, and the company's market value has fallen from a peak near $128 billion to roughly $75 billion. Robert Lea, an intelligence analyst, told Bloomberg: "This firm remains on a completely unsustainable commercial footing. Rising agentic AI will drive Z.ai's inference costs and losses higher."
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
The outside cybersecurity assessment available covers GLM-5.2 rather than GLM-5.3. In July, the UK AI Security Institute rated GLM-5.2 the strongest open-weight model it had tested for cybersecurity, comparable to closed models released four to seven months earlier. Through much of 2025, that gap had been six to ten months.
Hugging Face used GLM-5.2 to investigate a breach of its servers after guardrails on American frontier models declined to help.
The weights are due around August 28, after safety evaluation and hardening, marking the first time Z.ai has delayed a GLM weight release. Until then, access runs through the paid GLM Coding Plan, ZCode and controlled environments for selected security partners, with the most sensitive functions reserved for a trusted-access tier. Z.ai acknowledged that once the weights are public, it will no longer be able to control how people modify or use the model.
Frequently Asked Questions
What did GLM-5.3 score on CyberGym?
84.5%, up from GLM-5.2's 77.2%. Z.ai put Anthropic's Mythos 5 at 83.8% and OpenAI's GPT-5.6 Sol at 83.6% on the same test, which covered 1,507 tasks drawn from 188 software projects.
Why is Z.ai holding back the weights?
The company cited safety evaluation and hardening, and set the release for around August 28. It is the first time Z.ai has delayed a GLM weight release. Until then, access runs through the paid GLM Coding Plan, ZCode, and controlled environments for selected security partners.
Have the results been independently verified?
No. The cyber scores and the in-house Code Bench results were produced in Z.ai's own configuration, and no independent evaluator has replicated them. Artificial Analysis had not added the model as of August 14.
Does GLM-5.3 beat Anthropic on every cybersecurity benchmark?
No. On ExploitBench it scored 54.4% against 78.0% for Mythos 5, and on ExploitGym it completed 105 tasks under a two-hour budget against 181 for Mythos 5. The gap widens the further a benchmark moves up the exploitation chain.
What is in Z.ai's vulnerability ledger?
As of August 14 it listed 2,436 findings across 269 open-source projects, including 107 critical flaws and 990 rated high. Affected software included the Linux kernel, Redis, WebKit and FreeBSD. The oldest flaw was introduced in 1981.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR