Z.ai's GLM-5.3, built entirely through post-training on the same base model as its predecessor, more than doubled its exploit-chain scores while closing the coding gap with Anthropic's Fable 5.
Z.ai's GLM-5.3, released Aug. 14, derives every capability gain from scaled-up post-training on the same base model as GLM-5.2, yet more than doubled its exploitation benchmark scores and narrowed the coding gap with Anthropic's Fable 5. The Beijing-based company, also known as Zhipu, said the model is the strongest open-weights system it has measured for coding, with a 50 percent improvement over GLM-5.2 on its in-house Z.ai Code Bench.
"The further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2," the company said in its release, describing cyber capability that "developed faster than we expected" as training scaled. Z.ai said the model began reasoning across multiple stages of exploitation, forming coherent plans for complete chains rather than isolated bug-finding.
On ExploitBench, which requires deeper reasoning about real vulnerabilities, GLM-5.3 reached 54.4 percent, more than double GLM-5.2's 24.4 percent. On ExploitGym, it completed 105 tasks within two hours and 130 within six, versus 29 and 39 for its predecessor, with budgets normalized by per-model throughput. On CyberGym, it scored 84.5 percent, ahead of Mythos 5 at 83.8 percent and GPT-5.6 Sol at 83.6 percent. Coding gains were steepest on the longest-horizon evaluations: Terminal-Bench 3.0 jumped to 28.3 from 4.6, and DeepSWE v1.1 rose to 66.9 from 46.2.
The model identified 2,436 vulnerabilities across 269 open-source projects, including 1,097 rated critical or high severity, with the oldest flaw dating to 1981 and an average discovery lag of 26.6 years. Z.ai will release the weights in two weeks after safety evaluation and hardening, and the API now requires thinking enabled across three effort levels — low, high, and max — a breaking change for applications that previously ran with thinking switched off.
Post-Training Scaling, Not Architecture
GLM-5.3 uses the same roughly 700-billion-parameter base model as GLM-5.2, with all gains coming from environment scaling rather than architectural change. Z.ai built pipelines that synthesize long-horizon task environments end to end, with research agents converting real work patterns into runnable tasks and a judge agent verifying each is solvable. Some environments represent several days of work for an experienced engineer, such as an ML infrastructure task where the model must diagnose bottlenecks across a training stack, implement optimizations, and deliver a measurable end-to-end speedup.
The training runs on slime, Z.ai's open-source post-training framework, which the company said improved end-to-end RL throughput by more than 2.3 times for long-horizon coding tasks through workload-aware scheduling and hierarchical caching. Training-rollout consistency was tightened to a 1e-7 logprob difference, a reduction of more than 99.99 percent versus prior setups.
Competitive Positioning and the Open-Weight Question
The results place GLM-5.3 against a crowded field of Chinese open-weight challengers — DeepSeek-V4 Pro, Moonshot's Kimi K3, and Alibaba's Qwen3.8-Max — as well as closed models from Anthropic and OpenAI. On Z.ai's private Code Bench, GLM-5.3 reached 31.4 percent at roughly 50,000 output tokens per task at High effort, surpassing Claude Opus 4.8 at 29.5 percent with 120,000 tokens, while remaining behind Fable 5 at 39.5 percent at Max effort. On public suites, GLM-5.3 still trails GPT-5.6 Sol and Fable 5 on harder coding evaluations including Terminal-Bench 3.0 and DeepSWE, and Mythos 5 remains well ahead on ExploitGym at 181 and 247 tasks.
The two-week gap between announcement and weight release is doing work in this launch: the model being hardened is one Z.ai itself describes as having developed offensive security capability faster than expected, with its largest gains on the exploitation end of the chain. The findings feed Z.ai's public Security Disclosure Ledger, with 53 issues publicly disclosed and 2,383 under embargo at launch, including a use-after-free in the Linux kernel and a WebKit memory-handling flaw affecting Apple Safari.
Z.AI shares (02513.HK) reversed lower after the announcement, falling 3.9 percent to HK$1,266 in afternoon trading on turnover of HK$14.1 billion, even as Morgan Stanley raised its price target to HK$1,700. Whether independent evaluators replicate the vendor-reported numbers — particularly the in-house Code Bench results and cyber scores run in Z.ai's own harness configurations — will determine how much of this launch is a genuine step for open-weights coding models. The weight release, expected around the end of August, is when that testing begins.
This article is for informational purposes only and does not constitute investment advice.