AI Policy Wiki
Dashboard

GLM-5.3

medium confidence · updated 2026-08-17

Z.ai model launched August 14, 2026 — a post-training-only revision of the GLM-5.2 base model whose cyber-capability scores exceeded the company's stated expectations, prompting the first delay of a GLM open-weight release announced on safety grounds.

GLM-5.3 is a large language model launched on August 14, 2026 by Z.ai (formerly Zhipu AI), the successor to GLM-5.2 in the company's GLM line. Z.ai states that it did not retrain the base model, attributing the release's gains entirely to post-training, and titled its announcement "GLM-5.3: Frontier Coding with Emergent Cyber Capabilities" (Source: z.ai).

The release is distinguished less by its coding results than by how the company handled them. Z.ai reported that the model's offensive-security capability grew further during post-training than it had planned for, and said it would hold the weights for roughly two weeks after launch pending safety evaluation and hardening — a departure from GLM-5.2, whose MIT-licensed weights reached Hugging Face within days of its subscription launch (Source: z.ai; Source: reuters.com).

FieldValue
Developer[[companies/zhipu-ai\Z.ai (Zhipu AI)]]
LaunchedAugust 14, 2026 (API and GLM Coding Plan)
Predecessor[[models/glm-5-2\GLM-5.2]] (June 16, 2026)
Base modelUnchanged from GLM-5.2; gains attributed to post-training
WeightsAnnounced for release approximately two weeks after launch, pending safety evaluation and hardening
Access at launchZ.ai API and GLM Coding Plan; most sensitive cyber functions restricted to vetted users
Safety caseNo system card or published safety case as of August 15, 2026

Coding and agentic results

Z.ai reports a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, and describes the model as the strongest open-weights model for coding, claiming open-source state of the art on Terminal Bench 3.0 and Agents' Last Exam (Source: z.ai). The company positions the release as evidence that post-training alone can move a frontier open-weight model materially, without the cost of a new pretraining run — a framing echoed in trade coverage of the launch (Source: marktechpost.com).

The generation-over-generation figures Z.ai published are:

BenchmarkGLM-5.3GLM-5.2
Terminal-Bench 3.028.34.6
DeepSWE v1.166.946.2
Agents' Last Exam28.523.8

The Terminal-Bench 3.0 movement is the largest of the three in proportional terms, from a GLM-5.2 baseline near the floor of the benchmark (Source: z.ai). As with the cyber results below, all figures are Z.ai's own.

Cyber capability

The results Z.ai foregrounds are on three cyber evaluations that measure successive stages of the offensive-security chain. On CyberGym, which starts from white-box source code and tests whether a model can identify and validate a vulnerability by triggering the fault, Z.ai reports 84.5% against GLM-5.2's 77.2%, describing it as the leading result on the benchmark. On ExploitBench, which requires root-cause reasoning and a working exploit, it reports 54.4% against GLM-5.2's 24.4%. On ExploitGym, which counts exploitation tasks completed under time-normalized budgets, it reports 105 tasks within two hours and 130 within six, against 29 and 39 for GLM-5.2 (Source: z.ai).

BenchmarkGLM-5.3GLM-5.2Kimi K3DeepSeek-V4 Pro-0813Qwen3.8-MaxOpus 4.8(see note)GPT-5.6 Sol
CyberGym84.577.280.083.378.578.183.883.6
ExploitGym (2h / 6h)105 / 13029 / 3936 / 7014 / 2680 / 120181 / 247216 / 293
ExploitBench54.424.432.228.840.078.076.5

All figures are Z.ai's own, run in Z.ai's harness (Source: z.ai). The unlabeled column reflects a discrepancy in the source: Z.ai's prose attributes those figures to Claude Mythos 5, and Reuters reported them that way, while the table in the same announcement heads the column "Fable 5 (w/ fallback)" (Source: z.ai; Source: reuters.com). Anthropic has published no CyberGym figure for Mythos 5 that would settle the reading.

The pattern across the three evaluations is that GLM-5.3 leads on finding flaws and trails materially on weaponizing them: it is ahead of both reported Western frontier figures on CyberGym while reaching roughly two-thirds of their ExploitBench score and roughly half their ExploitGym task count. Z.ai's disclosed evaluation settings run the model inside Claude Code 2.1.207 at maximum reasoning effort, without web tools, under unlimited per-task timeouts, scored single-run pass@1 across CyberGym's 1,507 tasks, with a domain allowlist applied to prevent the agent from retrieving answers (Source: z.ai). No independent evaluator had reproduced the figures as of August 15, 2026; the weights required to do so were not yet public.

Vulnerability discovery programme

Alongside the benchmark figures, Z.ai reported that it had used the model with security teams in China to search open-source software, and that after expert review and deduplication the work identified 2,436 vulnerabilities across 269 projects, of which 1,097 were medium-to-high severity. As of the August 14, 2026 announcement, 53 had been publicly disclosed and 2,383 remained under embargo. Z.ai reported that the oldest flaw found had been introduced in 1981, and that the average vulnerability had persisted 26.6 years before discovery (Source: z.ai).

The figures are the developer's own and the embargoed majority is not independently checkable; no third party had confirmed the counts or the severity distribution as of August 17, 2026.

Release handling

Z.ai tied the delayed weight release explicitly to the cyber results, saying the weights would follow once safety evaluation and hardening were complete (Source: z.ai). Reporting on the launch describes a layered approach alongside the delay: the most sensitive cybersecurity functions restricted to a vetted-access program for launch partners, and prompt-screening guardrails trained to refuse malicious requests while preserving the model's usefulness for authorized penetration testing and bug fixing (Source: timesofindia.indiatimes.com).

Gabriel Wagner, an AI governance researcher at the Beijing-based consultancy Concordia AI, said that to the best of his knowledge this was the first time a Chinese lab had publicly justified a delayed open weight release on safety grounds, and read it as a sign that open-weight risk-management practice in China was becoming more sophisticated (Source: reuters.com). The delay is bounded — roughly two weeks — and Z.ai has not published a system card, a safety case, or the evaluation protocol the hardening period is meant to satisfy.

The release sits against the recurring policy question of open-weight frontier release: once weights are public under a permissive license, safety training can be removed locally and the release cannot be recalled, so a pre-release hardening period is the only point at which a developer's own controls bind. It also arrives during the period in which Anthropic has restricted its comparable cyber-specialized model, Mythos, to vetted partners under Project Glasswing, making the two the closest available comparison of restricted-access and open-weight handling of similar capability.

Open questions

  • Whether the weights are released on the announced timetable, under what license, and whether the "most sensitive functions" restriction survives an open-weight release in any enforceable form.
  • Whether independent evaluators reproduce the CyberGym, ExploitBench and ExploitGym figures once weights are available.
  • Whether Z.ai publishes the evaluation results or protocol from the hardening period, and whether the delayed-release precedent extends to later GLM models or to other Chinese developers.

Relationships