GLM-5.3: Frontier Coding Claims, No Weights Yet (2026 Launch Guide)

GLM-5.3 launched August 14, 2026 with big coding claims and cyber capabilities — but no open weights and no API for two weeks. What shipped, what's verified, and whether to wait.

Quick answer. GLM-5.3 launched August 14, 2026 as Z.ai's new coding flagship — but with no open weights and no API at launch. Access is only through the GLM Coding Plan ($18–$168/month) for roughly two weeks while safety hardening completes. It claims large agentic-coding gains and headline-grabbing cyber capabilities, all currently vendor-run and unverified by any independent board.

Z.ai released GLM-5.3 on August 14, 2026, under the striking title "Frontier coding with emergent cyber capabilities." It is the fourth frontier release in nine days — after Meta's Muse Spark 1.2, xAI's Grok 4.6 and DeepSeek's V4-Pro GA — and it is the most unusual of the four, for reasons that have little to do with benchmarks.

The two things GLM-5.3 launched without

For a model family whose identity is built on open weights, the launch itself is the story.

No open weights. GLM-5.2 and every recent GLM release hit Hugging Face on day one under MIT. GLM-5.3 did not. Z.ai says weights arrive "in two weeks after launch, once safety evaluation and hardening are complete" — roughly August 28. As of launch day there is no zai-org/GLM-5.3 repository, and no licence has been stated (5.2 was MIT).

No API. Both docs sites say "API coming soon," the pricing table has no GLM-5.3 row, and it is not on OpenRouter. The only way to use GLM-5.3 today is the GLM Coding Plan.

The charitable reading is caution about shipping a model with deliberately trained offensive-security skills. The sceptical reading, voiced by an LLM researcher in the launch discussion, is that a two-week exclusivity window converts launch-day attention into Coding Plan subscriptions. Both can be true at once; the weights date will settle it.

How do you access GLM-5.3 right now?

Through the Coding Plan, which was re-architected alongside the launch into a credit system:

PlanMonthlyAnnual (−30%)
Lite$18$12.60/mo
Pro$80$56/mo
Max$168$117.60/mo

The details that determine what you actually get:

  • Credit multipliers: GLM-5.3 consumes credits at 6.9x per input token, 1.7x cached, 24x per output token — so output-heavy agentic work drains quota far faster than the plan price suggests.
  • Off-peak is half price, and "peak" is narrowly defined: 14:00–18:00 UTC+8, Monday to Friday. Most of the world's working day is off-peak.
  • Silent routing: GLM-5.2 and 5.1 requests through the plan are now auto-routed to 5.3. If you pinned 5.2 for stability, you are on a new model as of today.
  • The plan exposes Anthropic- and OpenAI-compatible endpoints, so it plugs into Claude Code, Cline and similar tools. For the 1M context in Claude Code, the model string needs the suffix: glm-5.3[1m].

What are the specs?

SpecGLM-5.3
Context1M tokens
Max output128K
ModalityText only
ThinkingAlways on — low / high / max (default max); disabling now errors
Base modelSame as GLM-5.2 (~753B total, 256 experts / 8 active, per 5.2's config)
Knowledge cutoffNot published

The base being unchanged from 5.2 matters for local hosting: when weights do land, expect the same footprint — the smallest community GGUF of 5.2 is 217 GB, and the local-hosting community's measured verdict on consumer hardware was fractions of a token per second. This will not be a laptop model.

Do the benchmark claims hold up?

Nothing can be independently verified yet — GLM-5.3 is on zero external boards, and structurally cannot be until an API exists. What we can do is grade the provenance of Z.ai's own numbers, and it splits cleanly into good and bad.

The good

On Terminal-Bench 3.0, Z.ai's comparator numbers match the official leaderboard exactly — GLM-5.2 at 4.6, Claude Opus 4.8 at 21.1, GPT-5.6 Sol at 34.6, all under the same published harness. Quoting rivals' verified numbers accurately is better behaviour than most launches this month, and the claimed jump from 4.6 to 28.3 for GLM-5.3, if it verifies, is a genuine generational leap.

The bad

  • Terminal-Bench 2.1 is entirely self-run and reads hot — Z.ai cites Opus 4.8 at 85.0 where the official board says 78.9, and its own claimed 88.2 would beat the board's current #1. The only board-verified GLM number ever recorded is GLM-5.1 at 58.7%.
  • Claude Opus 5 is omitted from every row. The comparison set uses Opus 4.8 — Opus 5, the actual leader on the boards Z.ai cites, is absent. So is Grok 4.6.
  • Anti-cheat checks were removed on two benchmarks for Z.ai's own runs (disclosed, with LLM inspection substituted), and one benchmark's time budgets were rescaled by tokens-per-second in GLM's favour — roughly 2.9x the token budget of some rivals in the same nominal window.
  • The blog calls the same score "Fable 5" in the table and "Mythos 5" in the prose — an unexplained inconsistency that does not inspire confidence in editing rigour.

Our standing advice after a fortnight of these launches applies with extra force here: treat all of it as provisional until the API exists and a neutral harness runs it. Every vendor this month — Meta, xAI, Alibaba — published agentic numbers that shrank under independent measurement; GLM 5.2's own Terminal-Bench claim shrank by 13 points. See our guide to reading launch benchmarks.

What about the cyber capabilities?

This is the headline and the genuinely differentiated part — Z.ai deliberately trained on vulnerability-discovery data and published a real-world claim: 2,436 vulnerabilities found across 269 projects, with a public ledger.

We verified samples against MITRE and found genuine third-party corroboration — including a FreeBSD CVE crediting researchers "using GLM-5.1 from Z.ai" and a Red Hat advisory thanking Z.ai Security. The claim is real, though the count is cumulative since GLM-5.2 and only 53 entries are public. We cover the full picture — what is claimed, what verifies, what the security community makes of it — in our GLM-5.3 cyber capabilities deep-dive.

What is the early verdict from people who ran it?

Thin, and worth stating plainly: across roughly 770 comments and 2,300 combined upvotes on launch day, exactly four first-hand usage reports exist, and the only one with shared artifacts concluded Claude Opus 5 was "significantly better" on their task. There are no independent benchmark runs, no CTF reproductions, and no local runs — there cannot be, without weights.

The recurring complaints: the weights delay and scepticism of its safety rationale, the benchmark comparison-set choices, and the Coding Plan's output-token multiplier. The recurring praise: the price, the claimed Terminal-Bench 3.0 leap, and — notably — relief at a model that engages with security work rather than refusing it.

Should you use GLM-5.3?

Try it if you already have or want a GLM Coding Plan — $18/month against an effectively unlimited-feeling quota for off-peak hobby use remains one of the cheapest ways to run a large model inside Claude Code or Cline, and 5.3 is now what that plan serves.

Try it if you do security research — the deliberate cyber training and the CVE ledger are unique in the current field, and early users specifically valued that it does not refuse the work.

Wait if you need verified performance — no independent board lists it, and this month's pattern says launch numbers shrink.

Wait if you want open weights or API access — both are promised within roughly two weeks, and the licence question (MIT or something more restrictive, given the cyber framing) will be answered then. That answer matters more than any benchmark on the chart.

FAQ

When was GLM-5.3 released?

August 14, 2026, via the Z.ai blog. API and open weights were not included at launch.

Is GLM-5.3 open source?

Not yet. Z.ai says weights arrive roughly two weeks after launch, once safety evaluation completes. No licence has been stated; GLM-5.2 was MIT. This is a break from GLM's day-one open-weights pattern.

How can I use GLM-5.3 today?

Only through the GLM Coding Plan ($18/$80/$168 per month), which exposes Anthropic- and OpenAI-compatible endpoints for tools like Claude Code and Cline. Use glm-5.3[1m] for the 1M context in Claude Code.

What is GLM-5.3's context window?

1M tokens, with a 128K maximum output. It is text-only, with thinking always enabled at low/high/max effort.

Is GLM-5.3 better than GLM 5.2?

Z.ai claims large gains — most notably Terminal-Bench 3.0 rising from 4.6 to 28.3. The comparator numbers it quotes are accurate against official boards, but its own scores are self-run and unverified. The base model is unchanged from 5.2; the gains come from post-training.

Is GLM-5.3 better than Claude Opus 5?

Unknown — Z.ai omitted Opus 5 from every comparison row, benchmarking against Opus 4.8 instead. Opus 5 leads the official boards Z.ai cites, and the only detailed first-hand comparison so far favoured Opus 5.

Can I run GLM-5.3 locally?

Not yet, and realistically not on consumer hardware even when weights land — the base model is ~753B parameters and GLM-5.2's smallest community quantization is 217 GB.

Why does GLM-5.3 route my GLM-5.2 requests?

Z.ai now auto-routes Coding Plan requests for 5.2 and 5.1 to 5.3. There is no opt-out documented; if you pinned an older version for stability, you are on 5.3 as of August 14.