GLM-5.3 logo

GLM-5.3

Visit

Zhipu AI's newest flagship built for coding and cyber defense, squeezing 50% more coding capability out of the GLM-5 base through post-training scaling, ranked #1 among open-source models on Terminal Bench 3.0 and Agents' Last Exam (CLI).

Share:

GLM-5.3 is Zhipu AI's flagship coding and agentic model, released August 14, 2026. Zhipu kept the same base model as GLM-5.2 and pushed everything through extreme post-training scaling: tens of times more long-horizon task environments, richer environment types, and longer reinforcement learning on the IndexShare, SAO, and Slime framework stack. The result is a 50% jump in coding capability on Zhipu's internal evaluations, first place among open-source models on Terminal Bench 3.0 and Agents' Last Exam (CLI), and coding plus agent performance that Zhipu says approaches Claude Fable 5. Model weights are scheduled to go open source two weeks after release.

Core Features

  • Post-training scaling, same base: All gains come from the post-training stage on the unchanged GLM-5 MoE base, not from a bigger pre-training run.
  • Frontier-adjacent coding: Coding and agentic capability approaching Claude Fable 5, and the strongest open-weight coding experience among Chinese models, per Zhipu.
  • Cyber defense built in: Zhipu says GLM-5.3 ships with its most robust risk review system to date, aimed at making security capabilities broadly accessible.
  • Open weights in two weeks: MIT-style open release follows safety evaluations, continuing Zhipu's consistent open-source strategy.
  • Coding Plan integration: GLM Coding Plan quotas for all users were reset the same day, with the model consuming quota on the existing Lite/Pro/Max/Team subscription tiers.

Model Specifications

Specification GLM-5.3
Base model Same MoE base as GLM-5.2 (~744B, ~40B active)
Improvement source Post-training scaling only (RL on IndexShare/SAO/Slime)
Coding gain (internal) +50% vs GLM-5.2
Open weights Planned, ~2 weeks after release
Context Inherits the GLM-5 line's 1M-token window (per Z.ai docs)

Benchmark Highlights

  • Terminal Bench 3.0: 28.3 vs GLM-5.2's 4.6, roughly a 6x jump; #1 among open-source models.
  • DeepSWE v1.1: 66.9 vs 46.2 for GLM-5.2, a 44.8% gain on long-horizon software engineering.
  • Agents' Last Exam (CLI): 28.5 vs 23.8; top open-source score on cross-tool, long-horizon tasks.
  • AutomationBench: 48.2, beating Kimi K3 (46.7), Claude Fable 5 (46.2), and GPT-5.6 Sol.
  • CyberGym: 84.5%, above Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%).
  • ExploitGym: 105 tasks in 2 hours and 130 in 6 hours, versus 29 and 39 for GLM-5.2.
  • GDPval-AA v2: 1,769 points across 44 occupations, signaling coding gains extending into professional knowledge work.

Pricing

Per-token API pricing for GLM-5.3 had not been re-published at launch; it is expected to follow the GLM-5.2 structure (about $1.40 per 1M input tokens and $4.40 per 1M output tokens). GLM Coding Plan subscriptions (Lite/Pro/Max/Team) start around $18/month, and the model inherits the plan's peak/off-peak quota multipliers.

Use Cases

  • Agentic coding: Multi-step, long-horizon tasks where terminal and CLI tool use must stay coherent over a full session.
  • Cybersecurity work: Vulnerability discovery and verification from source code, with benchmark results closing the gap to closed frontier models.
  • Cost-sensitive teams: Frontier-adjacent coding at open-weight economics once weights ship.
  • Self-hosted deployments: MIT-licensed weights planned for vLLM, SGLang, and transformers stacks.

Advantages

  1. Post-training efficiency: Zhipu demonstrates that major capability gains are still available from RL and task-environment scaling on an existing base, without a new pre-training run.
  2. Open-source benchmark leader: First among open models on two headline coding/agent benchmarks the day it shipped.
  3. Security with openness: Cyber defense capability is being open-sourced rather than kept exclusive.

Tips

  1. Watch the weight release: The two-week open-source window is the event that decides whether GLM-5.3 becomes the default open coding model for agent teams.
  2. Benchmark the same session shape you run in production: Gains vary by benchmark, so test on your own long-horizon workflows before migrating from GLM-5.2.
  3. Reset your quota expectations: Coding Plan quotas were refreshed on release day; check "Usage Statistics" before heavy runs.

Conclusion

GLM-5.3 shows what post-training scaling can do: a 50% coding improvement over GLM-5.2 at the same base model, top open-source benchmark scores, and a concrete cyber-defense story. If open-weight coding performance near the frontier matters to you, GLM-5.3 is the release to watch this month.

Comments

No comments yet. Be the first to comment!