Tag

GLM

A collection of 14 posts

GLM

How to Run GLM-5.3-Flash Locally: Hardware & Setup

GLM-5.3-Flash ships 320B parameters in a 328 GB native-FP8 checkpoint, so the 18B active count tells you nothing about the memory you need. Here are the real hardware requirements by quantisation, plus working vLLM, SGLang, llama.cpp and Apple Silicon setups.

· 13 min read
GLM

GLM-5.3-Flash: Specs, Pricing and MIT Open Weights

Z.ai's GLM-5.3-Flash is the model that ran anonymously as Ox Alpha: a 320B-total, 18B-active natively multimodal MoE with a 1M-token context and MIT open weights. Here are the verified specs, pricing and benchmarks.

· 11 min read
AI

GLM-5.3 Weights Are Out: Full vs Flash (2026 Guide)

GLM-5.3 launched August 14, 2026 with big coding claims and cyber capabilities. The API and independent benchmarks have since arrived — the weights haven't. What shipped, what's verified, and whether to wait.

· 10 min read
Open Source LLMs

Ornith 1.0 vs GLM 5.2: Best Open Coding Model in 2026?

Two new MIT open-weights coding models shipped a day apart in June 2026. We compare architecture, coding benchmarks, local hardware, and API pricing for Ornith 1.0 vs GLM 5.2 — with an honest, no-hype verdict on which to pick.

· 15 min read
AI

How to Run GLM-5.2 Locally — and When to Pick 5.3-Flash

A practical walkthrough for self-hosting GLM-5.2 (744B MoE, 40B active) on llama.cpp. Quant tables, four hardware paths, exact install commands, verification, and a fallback to the Z.ai cloud API if your rig falls short.

· 14 min read
AI

GLM-5.2: 744B MoE, 1M Context, MIT-Licensed (2026)

Z.ai's GLM-5.2: 744B params (40B active), 1M-token context, MIT-licensed weights — still the newest GLM you can self-host while GLM-5.3's weights are pending. Architecture, benchmarks, pricing, and a 3-path local-inference playbook.

· 16 min read
AI

GLM 5.2 vs DeepSeek V4: The Open-Weights Coding Showdown (2026)

Two open-weights heavyweights from China go head-to-head for the agentic-coding throne. GLM 5.2 leads on context window; DeepSeek V4 leads on token economics. We break down cost, agentic strength, self-host paths, and pick a winner per workload.

· 13 min read