Tag

Open Source LLMs

A collection of 28 posts

GLM

How to Run GLM-5.3-Flash Locally: Hardware & Setup

GLM-5.3-Flash ships 320B parameters in a 328 GB native-FP8 checkpoint, so the 18B active count tells you nothing about the memory you need. Here are the real hardware requirements by quantisation, plus working vLLM, SGLang, llama.cpp and Apple Silicon setups.

· 13 min read
GLM

GLM-5.3-Flash: Specs, Pricing and MIT Open Weights

Z.ai's GLM-5.3-Flash is the model that ran anonymously as Ox Alpha: a 320B-total, 18B-active natively multimodal MoE with a 1M-token context and MIT open weights. Here are the verified specs, pricing and benchmarks.

· 11 min read
AI Models

Ox Alpha Was GLM-5.3-Flash: Specs, Price, Access

Ox Alpha was Z.ai's GLM-5.3-Flash, confirmed on 26 August 2026. The free stealth preview has ended and the OpenRouter listing is gone. Verified specs, MIT-licensed weights, real pricing, and how to use the model now.

· 13 min read
Qwen

Qwen 3.8 vs Qwen 3.6: Same Architecture, +14 Points (2026)

The config.json diff between Qwen 3.8-27B and Qwen 3.6-27B is empty — same layers, same hidden size, same vocab. Every gain came from post-training. Here's what actually changed, what the upgrade costs you, and who should stay on 3.6.

· 7 min read
Qwen

Qwen 3.7 vs Kimi K2.7: Best Open Agentic Coder in 2026?

Qwen3.7-Max is the closed flagship with higher vendor benchmarks and a 1M context; Kimi K2.7 Code is the open-weights, cheaper agentic specialist. We compare access, benchmarks, pricing, local feasibility, and which to use for autonomous coding.

· 13 min read
Ornith

Ornith 1.0 vs Claude Opus 4.8 for Coding (2026)

Ornith 1.0 is a free, MIT-licensed, self-hostable coding model. Opus 4.8 is the closed frontier flagship. A benchmark-grounded, harness-honest comparison of where each wins on agentic coding in 2026.

· 13 min read
DeepSeek V4

Qwen 3.7 vs DeepSeek V4: Best Open Coding Model in 2026?

DeepSeek V4 and Qwen 3.7 post near-identical coding benchmarks, but only one is actually open. A specifics-first comparison of architecture, local-run feasibility, API pricing, and license for developers choosing a coding model in 2026.

· 13 min read
Qwen

Qwen 3.6 27B as a Local Claude Code Replacement

A realistic, no-hype guide to running Qwen 3.6 27B locally as a Claude Code alternative: benchmarks vs Opus, the hardware and quant you actually need, how to wire it in, where it holds up, and what a hybrid setup actually saves you.

· 16 min read
Open Source LLMs

Ornith 1.0 vs GLM 5.2: Best Open Coding Model in 2026?

Two new MIT open-weights coding models shipped a day apart in June 2026. We compare architecture, coding benchmarks, local hardware, and API pricing for Ornith 1.0 vs GLM 5.2 — with an honest, no-hype verdict on which to pick.

· 15 min read
Gemma 4

Gemma 4 vs Qwen 3.5: Open LLM Comparison (2026)

A practical, size-tier-by-tier comparison of Google's Gemma 4 and Alibaba's Qwen 3.5 — benchmarks, coding, reasoning, multilingual, and how to run each locally in 2026.

· 8 min read