GLM-5.3 Weights Are Out: Full vs Flash (2026 Guide)

Quick answer. Both GLM-5.3 variants now have open weights. GLM-5.3-Flash (320B total / 18B active, natively multimodal) shipped August 26, 2026 under plain MIT. The full GLM-5.3 (753B) shipped August 28 under a custom licence that adds a security review for Model-as-a-Service operators above $10B revenue. Flash also spent six days on OpenRouter as the anonymous "Ox Alpha".

Z.ai released GLM-5.3 on August 14, 2026 under the title "Frontier coding with emergent cyber capabilities" — and then did something the GLM family had never done: it withheld the weights. For two weeks the interesting question about GLM-5.3 was not how it benchmarked but whether an open-weights lab would actually reopen a model it had described as dangerous.

That question is now answered, and the answer is more interesting than a simple yes. Z.ai shipped two models under two different licences, one of them via a week-long anonymous stealth run that briefly made it the most-used model on OpenRouter. This guide is the current state of the whole GLM-5.3 family, verified against Hugging Face, Z.ai's docs and OpenRouter on August 31, 2026.

Are GLM-5.3 weights available?

Yes — for both variants, as of the end of August 2026. The authoritative check is the zai-org organisation listing on Hugging Face, which currently carries four GLM-5.3 repositories:

  • zai-org/GLM-5.3-Flash — 321B params, native FP8, published August 26
  • zai-org/GLM-5.3-Flash-BF16 — the BF16 conversion, same date
  • zai-org/GLM-5.3 — 753B params, native FP8, published August 28
  • zai-org/GLM-5.3-BF16 — the BF16 conversion of the flagship

Z.ai's launch post promised weights "in two weeks after launch, once safety evaluation and hardening are complete." August 28 is exactly fourteen days after August 14. For once, a vendor's open-weights date landed on the day it said it would.

The catch is that the two releases do not carry the same licence, and most coverage has flattened that difference. Flash is MIT. The flagship is not.

What is the difference between GLM-5.3 and GLM-5.3-Flash?

They are not a big model and a distilled version of it. GLM-5.3 reuses GLM-5.2's 753B mixture-of-experts base and gets its gains from post-training. GLM-5.3-Flash is a newly trained base with a different architecture — and it is the first natively multimodal model in the GLM-5 family.

GLM-5.3GLM-5.3-Flash
ReleasedAug 14 (API) / Aug 28 (weights)Aug 26 (both)
Parameters753B total320B total / 18B active
ArchitectureSame MoE base as GLM-5.2New base; hybrid linear + sparse attention, 8 of 288 experts, 45 layers
Context1M tokens1M tokens (1,048,576 in config; OpenRouter lists 1,310,720)
Max output128K131K
ModalityText onlyText, image and video in; text out
LicenceCustom glm-5.3 licenceMIT, unmodified
WeightsYes — Aug 28, 2026Yes — Aug 26, 2026
API price / 1M$1.40 in · $0.26 cached · $4.40 out$0.15 in · $0.03 cached · $0.50 out (list)
FP8 weights on disk~756 GB (141 shards)~331 GB
AA Intelligence Index6057

The headline of that table is the last two rows read together. Flash gives up three points of measured intelligence and roughly nine-tenths of the API cost. Z.ai's own framing — "outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price" — is, for once, close to what the independent index shows: GLM-5.2 scored 53 on the same Artificial Analysis index, Flash scores 57.

Flash's price is currently even lower than the table suggests. Z.ai's pricing page is running a 50% launch discount through September 9, 2026, putting Flash at $0.075 / $0.015 / $0.25 per million. If you are evaluating it, evaluate it before that date.

We cover Flash's architecture and benchmarks in depth in the GLM-5.3-Flash complete guide, and the self-hosting path in how to run GLM-5.3-Flash locally.

Is GLM-5.3 open source, and what does the licence restrict?

GLM-5.3-Flash is straightforward: plain MIT, no field-of-use rider, no acceptable-use policy stapled on. Do what you like with it.

The flagship carries a bespoke document tagged glm-5.3 on the repo. Reading the licence file itself, it grants the usual rights — use, copy, modify, distribute, create derivatives, commercially — and then adds one substantive gate:

"If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue … exceeds 10 billion US dollars … in total over any consecutive 12 months, the Licensee must pass Z.AI's security review before using the Software."

"Model as a Service" is defined as giving third parties access to inference or fine-tuning where that third party meaningfully controls inputs, parameters or training data — so an API business, not a product with a model embedded in it. The review's scope is "determined by Z.AI at its discretion," with no published criteria, timeline or appeal route.

Practically: if you are a startup, an enterprise running this internally, or a researcher, this licence behaves exactly like MIT. If you are a hyperscaler planning to resell GLM-5.3 inference, you have an undefined approval step and a contact address (glmlicense@z.ai). Calling GLM-5.3 "MIT-licensed", as several outlets did in the first 48 hours, is wrong — that was Flash.

Was Ox Alpha GLM-5.3?

It was GLM-5.3-Flash, not the flagship — and the distinction matters, because the stealth run is the best independent evidence we have about Flash's real-world behaviour.

On August 20, 2026, a model with no listed maker appeared on OpenRouter as stealth/ox-alpha: 1,048,576-token context, 131K max output, text, image and video input, free for the duration. Over roughly six days it processed on the order of 23 trillion tokens — several times the next model on the platform — largely because it was a free frontier-class model with a million-token window and no rate card.

Z.ai confirmed on August 26, alongside the Flash launch, that Ox Alpha was GLM-5.3-Flash. OpenRouter's own listing for z-ai/glm-5.3-flash now states it directly.

What makes this more than trivia: a stealth window is an unusually clean read on a model, because nobody using it knew whose it was or what it was supposed to be good at. The usage numbers and the qualitative reaction from that week are the closest thing to an unbiased first impression any August release got. We break down the full timeline, the evidence trail that identified it, and what people found it good and bad at in our Ox Alpha stealth model guide.

What happened to the cyber-capability safety review?

Z.ai's stated reason for holding the weights was specific and unusual: the model had started reasoning across multiple stages of exploitation and assembling coherent end-to-end exploit chains — a capability the company says it did not set out to train. It reported 84.5 on CyberGym, ahead of Anthropic's Mythos 5, and a ledger of 2,436 vulnerabilities found across 269 projects.

The review concluded on schedule and the weights shipped on August 28. What Z.ai has not published is a hardening report: there is no public document describing what the two weeks of evaluation tested, what mitigations were applied, or what the residual risk assessment concluded. The delay was justified in public and resolved in silence.

That is worth naming plainly, because it sets a precedent. The open-weights community now has one instance of a lab pausing a release on cyber-capability grounds and then releasing anyway with no accompanying evidence about what changed. Whether the pause was a genuine safety process or a two-week Coding Plan exclusivity window is not something the public record can currently settle.

The underlying vulnerability claims, by contrast, do partly verify. We checked samples against MITRE and found genuine third-party corroboration, including a FreeBSD CVE crediting researchers using a GLM model and a Red Hat advisory thanking Z.ai Security. The full picture — what is claimed, what verifies, and what the security community makes of it — is in our GLM-5.3 cyber capabilities deep-dive.

How much does GLM-5.3 cost?

Three routes, and they price very differently.

Metered API. GLM-5.3 is $1.40 per million input, $0.26 cached, $4.40 output — identical to GLM-5.2's rate, so the upgrade is free in per-token terms. Flash is $0.15 / $0.03 / $0.50 at list, halved through September 9. OpenRouter mirrors both.

GLM Coding Plan. Still the cheapest way to run GLM inside Claude Code, Cline or similar, at $18 / $80 / $168 per month for Lite / Pro / Max. It runs on credits with two rolling caps — a 5-hour window and a 7-day allowance:

PlanMonthly5-hour creditsWeekly credits
Lite$182,00010,000
Pro$8012,00060,000
Max$16828,000140,000

The multipliers are where the real cost lives. Per 10,000 tokens, GLM-5.3 charges 6.9× input, 1.7× cached, 24× output; GLM-5.3-Flash charges 2.3× input, 0.56× cached, 8× output. Flash is roughly a third of the credit burn for the same work, which is what Z.ai means when it advertises "3× the usable quota". Output-heavy agentic loops drain a plan far faster than the monthly price implies, in both cases. Off-peak usage is half price, and peak is narrowly defined as Monday–Friday 14:00–18:00 SGT — most of the world's working day is off-peak.

One thing to know if you pinned an older model: Coding Plan requests for GLM-5.2, GLM-5.1 and GLM-4.7 are now auto-routed to the 5.3 family. If you were relying on GLM-5.2 for behavioural stability, you are not on it any more.

Self-hosting. Free per token, expensive in hardware — see below.

Do the benchmark claims hold up?

Better than most launches this month, with caveats that have not gone away.

The independent read comes from Artificial Analysis. GLM-5.3 scores 60 on its Intelligence Index at roughly $0.68 per index task, ranking #2 among open-weight models on that board. GLM-5.3-Flash scores 57 at around $0.09 per task, generating about 150M output tokens across the index run — noticeably more verbose than the ~110M median — at roughly 44 tokens/second. Both models buy their scores partly with output volume, which is exactly what makes the credit multipliers bite.

Z.ai's own numbers remain vendor-run and should be read as claims, not results:

  • GLM-5.3: Terminal-Bench 3.0 from 4.6 to 28.3 versus GLM-5.2, CyberGym 84.5, GDPval-AA 1769 Elo.
  • GLM-5.3-Flash: Terminal-Bench 2.1 at 84.3 against Claude Opus 4.8's 85.0 and GPT-5.6 Terra's 87.4; DeepSWE v1.1 at 63.4 versus GLM-5.2's 46.2; AutomationBench at 48.8 versus 26.2; Toolathlon Verified at 78.4.

The comparison-set problem from the launch post persists: Z.ai benchmarks against Claude Opus 4.8 rather than Opus 5, which leads the boards it cites. The cyber-specific benchmarks have still had no neutral reproduction. Our standing guidance on reading vendor launch tables — including which numbers historically shrink under independent measurement — is in our guide to reading launch benchmarks.

Can you run GLM-5.3 locally?

Now that weights exist, yes — on datacentre hardware. Neither variant is a laptop model.

GLM-5.3GLM-5.3-Flash
FP8 weights on disk~756 GB (141 shards)~331 GB
BF16 weights on disk~1.5 TB (282 shards)~640 GB
Minimum node8× H200 or 8× H208× H100/H200, or one GB200 tray at TP4
Full 1M context8× B200 (180 GB) with FP8 KV cacheHopper generation or newer required

Flash is the one worth attempting. At 18B active parameters it is dramatically cheaper to serve than its 320B total suggests, and a single 8×H100 node handles it. A100s and older will not work — the serving stack depends on FlashInfer kernels that need Hopper or newer.

The flagship at 753B is a genuine multi-node-budget proposition. For almost everyone, the API at $1.40/$4.40 is cheaper than the electricity, let alone the GPUs.

Which GLM-5.3 variant should you use?

A short decision rule, given everything above.

Use GLM-5.3-Flash for almost all agentic and coding work. It is three points behind the flagship on the only independent index either has, roughly a tenth of the API price, a third of the Coding Plan credit burn, genuinely MIT, natively multimodal, and self-hostable on a single node. The gap to the flagship is smaller than the price gap by an order of magnitude.

Use full GLM-5.3 when you need the extra headroom on hard reasoning, when you need text-only determinism from the mature GLM-5.2 base lineage, or when you are doing security research — the deliberate cyber training is specific to the flagship and is the one thing Flash does not replicate.

Check the licence first if you resell inference at scale. The $10B MaaS gate is narrow enough that it will never touch most readers, and absolute enough that it should be the first thing legal sees if it touches you.

And treat the cyber benchmarks as unsettled. The weights are out, which means neutral reproduction is finally possible; nobody has published one yet. If you are picking a model for security work on the strength of CyberGym 84.5, you are still trusting a vendor number.

FAQ

Are GLM-5.3 weights available?

Yes, both variants. GLM-5.3-Flash weights were published to Hugging Face on August 26, 2026 as zai-org/GLM-5.3-Flash, and the full 753B GLM-5.3 followed on August 28 as zai-org/GLM-5.3. BF16 conversions of both also exist. The two-week hold Z.ai announced at the August 14 launch was honoured to the day.

What is the difference between GLM-5.3 and GLM-5.3-Flash?

GLM-5.3 is 753B parameters, text-only, and reuses GLM-5.2's base with new post-training. GLM-5.3-Flash is a separate, newly trained 320B/18B-active model with hybrid linear-plus-sparse attention and native image and video input. Flash costs about a tenth as much per token and scores 57 to the flagship's 60 on the Artificial Analysis Intelligence Index.

Is GLM-5.3 open source?

Flash is — plain MIT with no restrictions. The full GLM-5.3 uses a custom glm-5.3 licence that permits commercial use, modification and redistribution but requires a Z.ai security review from Model-as-a-Service operators whose revenue exceeds $10B over any twelve months. For everyone below that line it behaves like MIT; it is not accurate to call the flagship MIT-licensed.

How much does GLM-5.3 cost?

The metered API is $1.40 per million input tokens, $0.26 cached and $4.40 output — the same rate as GLM-5.2. GLM-5.3-Flash lists at $0.15 / $0.03 / $0.50, halved to $0.075 / $0.015 / $0.25 during a launch promotion running to September 9, 2026. The GLM Coding Plan is $18, $80 or $168 per month and covers both models on a credit system.

Was Ox Alpha GLM-5.3?

Ox Alpha was GLM-5.3-Flash, not the flagship. It ran anonymously on OpenRouter as stealth/ox-alpha from August 20 to 26, 2026 with a 1M-token context and free access, processed roughly 23 trillion tokens in that window, and briefly topped the platform's usage chart. Z.ai confirmed its identity when it launched Flash on August 26.

Can I run GLM-5.3 locally?

Only on datacentre GPUs. The flagship's FP8 weights are about 756 GB and need at least an 8×H200 node, with 8×B200 required for the full 1M context. GLM-5.3-Flash is far more tractable at roughly 331 GB in FP8 on a single 8×H100 node, but still needs Hopper-generation hardware or newer. Neither runs on consumer machines.

Does GLM-5.3 replace GLM-5.2?

On the GLM Coding Plan, effectively yes — requests for GLM-5.2, GLM-5.1 and GLM-4.7 are auto-routed to the 5.3 family with no documented opt-out. On the metered API, GLM-5.2 is still listed at the same $1.40/$4.40 pricing, and its weights remain on Hugging Face under their original terms if you need to pin a specific version.