Z.ai GLM-5.3 release arrives with coding claims and a safety hold
Z.ai made GLM-5.3 available through its coding services, while delaying downloadable weights for cyber-safety testing.
By Theo Nakamura · Staff Writer
· 3 min read
The Z.ai GLM-5.3 release puts a new coding-focused AI model into the company’s hosted products, but not yet into the hands of developers as downloadable weights. Z.ai announced the model on Aug. 14 and made it available through its Coding Plan and ZCode, according to the company and reporting by Yahoo Tech. For investors and developers tracking the open-model market, the distinction matters: Z.ai’s claim of leadership applies to weights it said would arrive later, after safety work.
Z.ai said it expects to publish GLM-5.3’s weights two weeks after launch, once it completes safety evaluations and hardening. Downloadable weights would allow others to deploy the model themselves, whereas the initial release is access through Z.ai’s services. API access was also slated to follow the safety review, Yahoo Tech reported.
How does GLM-5.3 compare with other coding models?
Z.ai calls GLM-5.3 its strongest open-weight coding model, a vendor claim based on its published tests. The company said the model uses the same underlying base model as GLM-5.2, with improvements coming from post-training, or additional training after the base model was built. Z.ai said it added more task environments, a wider mix of tasks and more computing resources.
The reported gains over GLM-5.2 were substantial on several coding evaluations. Z.ai listed a 28.3 score on Terminal Bench 3.0, up from 4.6 for GLM-5.2, and a 66.9 score on DeepSWE v1.1, compared with 46.2 for its predecessor.
The comparisons also put limits on a broad leadership reading. In Z.ai’s table, Fable 5 scored 33.7 and GPT-5.6 Sol scored 34.6 on Terminal Bench 3.0, ahead of GLM-5.3. On DeepSWE v1.1, Kimi K3 scored 67.5, Fable 5 reached 69.7 and GPT-5.6 Sol posted 72.7, all above GLM-5.3’s 66.9.
On Z.ai Code Bench, its private coding evaluation, Z.ai reported GLM-5.3 completed 34.5% of tasks at its maximum effort setting while using about 75,000 output tokens per task. GLM-5.2 recorded 23.4% at roughly 96,000 tokens, according to the company. Z.ai said Fable 5 remained ahead at 39.5% under the same maximum-effort setting. Because Code Bench is an in-house test, these results have not been independently verified in the evidence available.
Why is Z.ai delaying GLM-5.3’s downloadable weights?
The delay centers on cybersecurity capability. Z.ai said GLM-5.3 scored 84.5% on CyberGym and that its ability to work through multi-stage exploitation developed faster than expected as post-training expanded. The company said it added vulnerability-discovery data and controlled cyber environments to the training mix.
Axios reported that Z.ai is using a tiered-access program for selected security partners in controlled settings while it tests and strengthens safety controls. Z.ai has presented the model as potentially useful to defenders, while Axios noted that cyber-capable AI can also be used by attackers.
The key near-term question is whether the promised weights arrive on Z.ai’s timetable and under what safeguards. Until then, GLM-5.3 is a hosted coding product with vendor-reported benchmark gains, rather than a publicly downloadable open-weight model.
This story draws on original reporting from Decrypt.