Overview
The GLM-5.3 line is Zhipu AI's current generation, and it arrived in two pieces eight days apart. GLM-5.3 landed on August 18, 2026 — the date on Zhipu's own release notes, which credit it with "a 50% gain over GLM-5.2 on Z.ai Code Bench" and "emergent cybersecurity capabilities." GLM-5.3-Flash followed on August 26, 2026, described by the same notes as combining "native visual capabilities" with an "efficient hybrid architecture" of linear and sparse attention.
The two are easy to misread as a big model and its small sibling. They are not. GLM-5.3 reuses the GLM-5.2 base model unchanged and takes its gains entirely from post-training — a cheap, fast way to ship capability, and one that leaves the underlying serving cost where it was. GLM-5.3-Flash is a new architecture: 320B total parameters with 18B activated per token, natively multimodal, and built specifically so that a million-token context stops being expensive.
Their licensing postures also diverge, which matters more than the version numbers suggest. GLM-5.3-Flash shipped with MIT weights on Hugging Face on launch day. GLM-5.3's weights, as of August 27, 2026, have not been published and no license has been stated — Zhipu signalled a release roughly two weeks after launch, pending safety evaluation. For now the flagship is API-only and the efficiency model is the open one.
Capabilities
GLM-5.3 — coding and security, via post-training
- Coding: 34.5% on Z.ai Code Bench at
maxeffort, against GLM-5.2's 23.4% — and it spends fewer tokens getting there, roughly 75K against 96K. - Vulnerability discovery: 84.5% on CyberGym, up from 77.2%. Zhipu characterises the security jump as emergent rather than targeted.
- Exploit construction: 54.4% on ExploitBench, more than double GLM-5.2's 24.4%.
- Mandatory reasoning: three effort levels —
low,high,max. Reasoning can no longer be switched off.
GLM-5.3-Flash — multimodal agents on a long, cheap context
- Native vision: reads interfaces, rendering results and interaction feedback directly, alongside images, video and files. This is multimodal input aimed at software that operates a screen, not at captioning.
- Agentic coding: 63.4 on DeepSWE v1.1 against GLM-5.2's 46.2; 48.8 on AutomationBench against 26.2.
- Terminal work: 84.3 on Terminal-Bench 2.1, against 85.0 for Claude Opus 4.8 in Zhipu's own table.
- Independently scored: 57 on the Artificial Analysis Intelligence Index v4.1.1, ranking 3rd of 109 models measured.
- Cheap long context: 3.01x less attention computation and a 4.44x smaller KV cache than GLM-5.3.
Technical Specifications
GLM-5.3
- API model ID:
glm-5.3 - Base model: identical to GLM-5.2; improvements are post-training only
- Depth: 92 layers, per the comparison Zhipu publishes in the GLM-5.3-Flash documentation
- Context window: 1M tokens; max output 128K tokens
- Weights: not published as of August 27, 2026; license unstated
GLM-5.3-Flash
- API model ID:
glm-5.3-flash - Hugging Face repository:
zai-org/GLM-5.3-Flash, MIT - Parameters: 320B total, 18B activated per token
- Routing: 8 of 288 experts, Mixture-of-Experts
- Depth: 45 layers, interleaving KDA linear-attention layers with NoPE sparse MLA layers
- Extras: IndexPool key-vector compression; Manifold-Constrained Hyper-Connections (mHC) for scaling efficiency
- Weights format: FP8, roughly 306 GiB
- Training corpus: 30T tokens, multimodal
- Context window: 1M tokens
What the hybrid attention buys
A long context window is charged twice on a conventional transformer: attention compute that scales with sequence length, and a KV cache that grows linearly and then dominates GPU memory. GLM-5.3-Flash halves the depth relative to GLM-5.3 and splits the remaining layers between a linear-cost recurrence (KDA) and sparse full attention (NoPE MLA). Zhipu's measured result is 3.01x less attention compute and a 4.44x smaller KV cache. That is the whole product thesis: the same class of capability, priced so a million-token prompt is routine.
Conflicting context figures
Z.ai's documentation and release notes give 1M tokens (1,048,576) for GLM-5.3-Flash. The Hugging Face card states 300,000; OpenRouter advertises 1,310,720 with a 48,000-token completion cap. These describe different things — architectural support, host configuration, and card vintage — so confirm the ceiling on the endpoint you actually call rather than trusting the headline.
Use Cases
- Screen-driven agents and GUI automation: GLM-5.3-Flash reads rendering results and interaction feedback natively, which is what a computer-use loop needs and what a text-only model has to be bolted onto.
- Vision-to-code frontend work: mockups and screenshots in, runnable interface code out, without a separate captioning stage.
- Whole-repository sessions on a budget: the 4.44x KV-cache reduction is what makes holding a large codebase in context economically ordinary rather than a special occasion.
- Defensive security research: GLM-5.3's CyberGym and ExploitBench numbers point at vulnerability triage and patch validation — with the caveat that these figures are unreplicated.
- Office and document pipelines: PPTX, PDF, DOCX and XLSX deliverables, plus financial workflow and research tasks Zhipu calls out explicitly.
- Token-heavy batch work: at $0.15 input and $0.50 output, GLM-5.3-Flash is priced for jobs where volume matters more than latency.
Performance / Benchmarks
Vendor-reported, from Zhipu's documentation. Blank cells mean the comparison was not published.
| Benchmark | GLM-5.2 | GLM-5.3 | GLM-5.3-Flash | Claude Opus 4.8 |
|---|---|---|---|---|
| Z.ai Code Bench | 23.4 | 34.5 | 29.0 | 29.5 |
| CyberGym | 77.2 | 84.5 | — | — |
| ExploitBench | 24.4 | 54.4 | — | — |
| DeepSWE v1.1 | 46.2 | — | 63.4 | — |
| AutomationBench | 26.2 | — | 48.8 | — |
| Terminal-Bench 2.1 | 81.0 | — | 84.3 | 85.0 |
Two things are worth separating here. The GLM-5.2 columns are the vendor grading its own predecessor, the easiest comparison to arrange. The Opus 4.8 column is the one Zhipu chose to publish anyway, and it does not hide the remaining gap.
Independent measurement
Artificial Analysis has run GLM-5.3-Flash through its own harness — the first outside read on the model. It scores 57 on the Intelligence Index v4.1.1, 3rd of 109 models, across nine evaluations including Terminal-Bench 2.1, GPQA Diamond, Humanity's Last Exam, SciCode and AA-LCR.
The operational figures are less flattering and matter more inside an agent loop:
| Measure | GLM-5.3-Flash | Median of measured models |
|---|---|---|
| Output speed | 50.2 tokens/s | 67 tokens/s |
| Time to first token | 1.47 s | — |
| Output tokens to complete the index | 150M | 100M |
| Cost to run the full index | $138.02 | — |
The model is cheap per token, slow per token, and roughly 50% more verbose than the median. For batch work the token price dominates and the trade is clearly good; for interactive latency it is closer than the rate card implies.
Limitations
- The flagship is not open: GLM-5.3 weights are unpublished and unlicensed as of August 27, 2026. Anything requiring on-premises deployment has to use GLM-5.3-Flash or stay on GLM-5.2.
- Reasoning cannot be disabled: every GLM-5.3 call pays some thinking cost, which puts a floor under latency and output tokens for trivial requests.
- GLM-5.3-Flash is slow and verbose: 50.2 tokens/s against a 67 median, and 50% more output tokens to finish the same evaluation set.
- Self-hosting GLM-5.3-Flash is still a cluster job: ~306 GiB of FP8 weights is a multi-GPU deployment, notwithstanding the 18B activated slice.
- No activated-parameter count for GLM-5.3: serving cost and latency remain hard to model in advance.
- Benchmarks are largely vendor-reported: Artificial Analysis covers GLM-5.3-Flash; the CyberGym, ExploitBench and Code Bench figures have no independent replication.
- Two release dates in circulation: several write-ups date GLM-5.3 to August 14, 2026. Zhipu's dated release note says August 18, and that is the figure this page uses.
- Conflicting context ceilings: 1M, 300K and 1.31M appear across the docs, the model card and OpenRouter.
Pricing & Access
From the z.ai pricing page, per 1M tokens in USD.
| Model | Input | Cached input | Output |
|---|---|---|---|
| GLM-5.3 | $1.40 | $0.26 | $4.40 |
| GLM-5.3-Flash | $0.15 | $0.03 | $0.50 |
| GLM-5.2 | $1.40 | $0.26 | $4.40 |
| GLM-5.1 | $1.40 | $0.26 | $4.40 |
| GLM-5 | $1.00 | $0.20 | $3.20 |
| GLM-4.7 | $0.60 | $0.11 | $2.20 |
| GLM-4.7-Flash | Free | Free | Free |
A promotion halves the GLM-5.3-Flash rates to $0.075 input and $0.25 output through September 9, 2026. Cached input storage is listed as free for a limited time across the line.
Access options:
- z.ai API — the primary managed endpoint for both models
- GLM Coding Plan — GLM-5.3-Flash is fully available on every tier, with 3x the quota of GLM-5.3
- Hugging Face — MIT weights at
zai-org/GLM-5.3-Flash; GLM-5.2 remains atzai-org/GLM-5.2 - OpenRouter — third-party routing, where the model ran as "ox-alpha" before launch
- vLLM / SGLang — high-performance local serving
Ecosystem & Tools
- GLM-5.3 documentation — reasoning effort levels, function calling, context caching, structured output, MCP
- GLM-5.3-Flash documentation — multimodal input, architecture and efficiency figures
- GLM-5.3-Flash on Hugging Face — MIT weights and the benchmark table
- z.ai — platform, console and chat interface
- vLLM — production serving with tensor and pipeline parallelism
- SGLang — structured generation and high-throughput serving
- MCP — native Model Context Protocol tool integration
Community & Resources
- z.ai release notes — dated entries for GLM-5.3 (August 18, 2026) and GLM-5.3-Flash (August 26, 2026)
- GLM-5.3-Flash on Artificial Analysis — independent Intelligence Index v4.1.1 score, speed and cost measurements
- Pricing
- zai-org on Hugging Face
- Our launch coverage: GLM-5.3-Flash and GLM-5
- Compare with DeepSeek V4, Kimi K2.6, Tencent Hy3, and Ling-2.6-1T