Developer
Zhipu AI

GLM-5.3

Zhipu AI's current line: GLM-5.3, a post-trained coding and security flagship, and GLM-5.3-Flash, a 320B/18B multimodal MoE with hybrid attention.

Updated

Released
Aug 18, 2026
Type
Language Model
Context window
1M tokens
Pricing
$1.40 / $4.40 per Mtok
License
MIT for GLM-5.3-Flash; GLM-5.3 weights not yet released
On this page

Overview

The GLM-5.3 line is Zhipu AI's current generation, and it arrived in two pieces eight days apart. GLM-5.3 landed on August 18, 2026 — the date on Zhipu's own release notes, which credit it with "a 50% gain over GLM-5.2 on Z.ai Code Bench" and "emergent cybersecurity capabilities." GLM-5.3-Flash followed on August 26, 2026, described by the same notes as combining "native visual capabilities" with an "efficient hybrid architecture" of linear and sparse attention.

The two are easy to misread as a big model and its small sibling. They are not. GLM-5.3 reuses the GLM-5.2 base model unchanged and takes its gains entirely from post-training — a cheap, fast way to ship capability, and one that leaves the underlying serving cost where it was. GLM-5.3-Flash is a new architecture: 320B total parameters with 18B activated per token, natively multimodal, and built specifically so that a million-token context stops being expensive.

Their licensing postures also diverge, which matters more than the version numbers suggest. GLM-5.3-Flash shipped with MIT weights on Hugging Face on launch day. GLM-5.3's weights, as of August 27, 2026, have not been published and no license has been stated — Zhipu signalled a release roughly two weeks after launch, pending safety evaluation. For now the flagship is API-only and the efficiency model is the open one.

Capabilities

GLM-5.3 — coding and security, via post-training

  • Coding: 34.5% on Z.ai Code Bench at max effort, against GLM-5.2's 23.4% — and it spends fewer tokens getting there, roughly 75K against 96K.
  • Vulnerability discovery: 84.5% on CyberGym, up from 77.2%. Zhipu characterises the security jump as emergent rather than targeted.
  • Exploit construction: 54.4% on ExploitBench, more than double GLM-5.2's 24.4%.
  • Mandatory reasoning: three effort levels — low, high, max. Reasoning can no longer be switched off.

GLM-5.3-Flash — multimodal agents on a long, cheap context

  • Native vision: reads interfaces, rendering results and interaction feedback directly, alongside images, video and files. This is multimodal input aimed at software that operates a screen, not at captioning.
  • Agentic coding: 63.4 on DeepSWE v1.1 against GLM-5.2's 46.2; 48.8 on AutomationBench against 26.2.
  • Terminal work: 84.3 on Terminal-Bench 2.1, against 85.0 for Claude Opus 4.8 in Zhipu's own table.
  • Independently scored: 57 on the Artificial Analysis Intelligence Index v4.1.1, ranking 3rd of 109 models measured.
  • Cheap long context: 3.01x less attention computation and a 4.44x smaller KV cache than GLM-5.3.

Technical Specifications

GLM-5.3

  • API model ID: glm-5.3
  • Base model: identical to GLM-5.2; improvements are post-training only
  • Depth: 92 layers, per the comparison Zhipu publishes in the GLM-5.3-Flash documentation
  • Context window: 1M tokens; max output 128K tokens
  • Weights: not published as of August 27, 2026; license unstated

GLM-5.3-Flash

  • API model ID: glm-5.3-flash
  • Hugging Face repository: zai-org/GLM-5.3-Flash, MIT
  • Parameters: 320B total, 18B activated per token
  • Routing: 8 of 288 experts, Mixture-of-Experts
  • Depth: 45 layers, interleaving KDA linear-attention layers with NoPE sparse MLA layers
  • Extras: IndexPool key-vector compression; Manifold-Constrained Hyper-Connections (mHC) for scaling efficiency
  • Weights format: FP8, roughly 306 GiB
  • Training corpus: 30T tokens, multimodal
  • Context window: 1M tokens

What the hybrid attention buys

A long context window is charged twice on a conventional transformer: attention compute that scales with sequence length, and a KV cache that grows linearly and then dominates GPU memory. GLM-5.3-Flash halves the depth relative to GLM-5.3 and splits the remaining layers between a linear-cost recurrence (KDA) and sparse full attention (NoPE MLA). Zhipu's measured result is 3.01x less attention compute and a 4.44x smaller KV cache. That is the whole product thesis: the same class of capability, priced so a million-token prompt is routine.

Conflicting context figures

Z.ai's documentation and release notes give 1M tokens (1,048,576) for GLM-5.3-Flash. The Hugging Face card states 300,000; OpenRouter advertises 1,310,720 with a 48,000-token completion cap. These describe different things — architectural support, host configuration, and card vintage — so confirm the ceiling on the endpoint you actually call rather than trusting the headline.

Use Cases

  • Screen-driven agents and GUI automation: GLM-5.3-Flash reads rendering results and interaction feedback natively, which is what a computer-use loop needs and what a text-only model has to be bolted onto.
  • Vision-to-code frontend work: mockups and screenshots in, runnable interface code out, without a separate captioning stage.
  • Whole-repository sessions on a budget: the 4.44x KV-cache reduction is what makes holding a large codebase in context economically ordinary rather than a special occasion.
  • Defensive security research: GLM-5.3's CyberGym and ExploitBench numbers point at vulnerability triage and patch validation — with the caveat that these figures are unreplicated.
  • Office and document pipelines: PPTX, PDF, DOCX and XLSX deliverables, plus financial workflow and research tasks Zhipu calls out explicitly.
  • Token-heavy batch work: at $0.15 input and $0.50 output, GLM-5.3-Flash is priced for jobs where volume matters more than latency.

Performance / Benchmarks

Vendor-reported, from Zhipu's documentation. Blank cells mean the comparison was not published.

BenchmarkGLM-5.2GLM-5.3GLM-5.3-FlashClaude Opus 4.8
Z.ai Code Bench23.434.529.029.5
CyberGym77.284.5——
ExploitBench24.454.4——
DeepSWE v1.146.2—63.4—
AutomationBench26.2—48.8—
Terminal-Bench 2.181.0—84.385.0

Two things are worth separating here. The GLM-5.2 columns are the vendor grading its own predecessor, the easiest comparison to arrange. The Opus 4.8 column is the one Zhipu chose to publish anyway, and it does not hide the remaining gap.

Independent measurement

Artificial Analysis has run GLM-5.3-Flash through its own harness — the first outside read on the model. It scores 57 on the Intelligence Index v4.1.1, 3rd of 109 models, across nine evaluations including Terminal-Bench 2.1, GPQA Diamond, Humanity's Last Exam, SciCode and AA-LCR.

The operational figures are less flattering and matter more inside an agent loop:

MeasureGLM-5.3-FlashMedian of measured models
Output speed50.2 tokens/s67 tokens/s
Time to first token1.47 s—
Output tokens to complete the index150M100M
Cost to run the full index$138.02—

The model is cheap per token, slow per token, and roughly 50% more verbose than the median. For batch work the token price dominates and the trade is clearly good; for interactive latency it is closer than the rate card implies.

Limitations

  • The flagship is not open: GLM-5.3 weights are unpublished and unlicensed as of August 27, 2026. Anything requiring on-premises deployment has to use GLM-5.3-Flash or stay on GLM-5.2.
  • Reasoning cannot be disabled: every GLM-5.3 call pays some thinking cost, which puts a floor under latency and output tokens for trivial requests.
  • GLM-5.3-Flash is slow and verbose: 50.2 tokens/s against a 67 median, and 50% more output tokens to finish the same evaluation set.
  • Self-hosting GLM-5.3-Flash is still a cluster job: ~306 GiB of FP8 weights is a multi-GPU deployment, notwithstanding the 18B activated slice.
  • No activated-parameter count for GLM-5.3: serving cost and latency remain hard to model in advance.
  • Benchmarks are largely vendor-reported: Artificial Analysis covers GLM-5.3-Flash; the CyberGym, ExploitBench and Code Bench figures have no independent replication.
  • Two release dates in circulation: several write-ups date GLM-5.3 to August 14, 2026. Zhipu's dated release note says August 18, and that is the figure this page uses.
  • Conflicting context ceilings: 1M, 300K and 1.31M appear across the docs, the model card and OpenRouter.

Pricing & Access

From the z.ai pricing page, per 1M tokens in USD.

ModelInputCached inputOutput
GLM-5.3$1.40$0.26$4.40
GLM-5.3-Flash$0.15$0.03$0.50
GLM-5.2$1.40$0.26$4.40
GLM-5.1$1.40$0.26$4.40
GLM-5$1.00$0.20$3.20
GLM-4.7$0.60$0.11$2.20
GLM-4.7-FlashFreeFreeFree

A promotion halves the GLM-5.3-Flash rates to $0.075 input and $0.25 output through September 9, 2026. Cached input storage is listed as free for a limited time across the line.

Access options:

  • z.ai API — the primary managed endpoint for both models
  • GLM Coding Plan — GLM-5.3-Flash is fully available on every tier, with 3x the quota of GLM-5.3
  • Hugging Face — MIT weights at zai-org/GLM-5.3-Flash; GLM-5.2 remains at zai-org/GLM-5.2
  • OpenRouter — third-party routing, where the model ran as "ox-alpha" before launch
  • vLLM / SGLang — high-performance local serving

Ecosystem & Tools

  • GLM-5.3 documentation — reasoning effort levels, function calling, context caching, structured output, MCP
  • GLM-5.3-Flash documentation — multimodal input, architecture and efficiency figures
  • GLM-5.3-Flash on Hugging Face — MIT weights and the benchmark table
  • z.ai — platform, console and chat interface
  • vLLM — production serving with tensor and pipeline parallelism
  • SGLang — structured generation and high-throughput serving
  • MCP — native Model Context Protocol tool integration

Community & Resources

Frequently Asked Questions

August 18, 2026, the date on Zhipu's own release notes entry. Several third-party write-ups circulate August 14; the vendor's dated release note is the primary source. GLM-5.3-Flash followed on August 26, 2026.
They are not the same model shrunk. GLM-5.3 reuses the GLM-5.2 base model and gains entirely from post-training. GLM-5.3-Flash is a distinct 320B/18B-active architecture, natively multimodal, with hybrid sparse and linear attention built for cheap long context.
320 billion total with 18 billion activated per token, routed through 8 of 288 experts across 45 layers. GLM-5.3 itself has no published parameter count; it shares the GLM-5.2 base.
Not yet. As of August 27, 2026 the GLM-5.3 weights have not appeared on Hugging Face and the license has not been stated. GLM-5.3-Flash is available now under MIT as zai-org/GLM-5.3-Flash.
No. Z.ai's documentation states that disabling reasoning is no longer supported. Effort is instead selected from three levels — low, high and max — so every call carries some thinking cost.
GLM-5.3 is $1.40 per 1M input tokens, $0.26 cached and $4.40 output — the same rate card as GLM-5.2. GLM-5.3-Flash lists at $0.15 input and $0.50 output, halved to $0.075 and $0.25 by a promotion running through September 9, 2026.
Z.ai reports 34.5% on its own Z.ai Code Bench at max effort using roughly 75K output tokens, against GLM-5.2's 23.4% at 96K tokens — a claimed 50% gain that also spends fewer tokens. The benchmark is the vendor's own.
Z.ai describes it as emergent. On CyberGym vulnerability discovery it reports 84.5% against GLM-5.2's 77.2%, and on ExploitBench 54.4% against 24.4%. These are vendor-reported and have not been independently replicated.
1M tokens for both models, with 128K maximum output for GLM-5.3. Host-configured limits vary — the GLM-5.3-Flash card on Hugging Face lists 300,000 and OpenRouter advertises 1,310,720 — so verify the endpoint you call.
Yes. It was served without attribution as 'ox-alpha' on OpenRouter and OpenCode before the August 26 reveal, where it became one of the most-used models of the week.

Explore More Models

Discover other AI models and compare their capabilities.