Qwen 3.8 Model Lineup: Every Variant and Which to Use
Every Qwen 3.8 variant in one table: 27B, Flash-Next, 2.4T-A95B and Max, with params, context, licence and price, plus a single rule for picking one.
A collection of 30 posts
Every Qwen 3.8 variant in one table: 27B, Flash-Next, 2.4T-A95B and Max, with params, context, licence and price, plus a single rule for picking one.
Alibaba's cheap multimodal tier: $0.15/$0.47 per 1M tokens, 1M context, and open weights under the Qwen Community License. Specs, benchmarks and how it compares to GLM-5.3-Flash and DeepSeek V4-Flash.
GLM-5.3-Flash is 320B-A18B, multimodal and roughly 9x cheaper than GLM-5.2's 744B-A40B. A verified head-to-head on specs, benchmarks, cost and self-hosting — plus the cases where GLM-5.2 still wins.
GLM-5.3-Flash ships 320B parameters in a 328 GB native-FP8 checkpoint, so the 18B active count tells you nothing about the memory you need. Here are the real hardware requirements by quantisation, plus working vLLM, SGLang, llama.cpp and Apple Silicon setups.
Z.ai's GLM-5.3-Flash is the model that ran anonymously as Ox Alpha: a 320B-total, 18B-active natively multimodal MoE with a 1M-token context and MIT open weights. Here are the verified specs, pricing and benchmarks.
Ox Alpha was Z.ai's GLM-5.3-Flash, confirmed on 26 August 2026. The free stealth preview has ended and the OpenRouter listing is gone. Verified specs, MIT-licensed weights, real pricing, and how to use the model now.
The 17GB Apache-2.0 model scoring 52 on Artificial Analysis: which benchmark numbers to trust, the overthinking fix, VRAM needs, and how it compares to Claude and the 2.4T flagship.
The config.json diff between Qwen 3.8-27B and Qwen 3.6-27B is empty — same layers, same hidden size, same vocab. Every gain came from post-training. Here's what actually changed, what the upgrade costs you, and who should stay on 3.6.
A practical, honest guide to running MiniMax M3 (428B MoE) locally: VRAM/RAM math, quant options, and the Ollama, vLLM, and LM Studio paths.
DiffusionGemma 26B-A4B is Google’s first open-weight text-diffusion LLM — a 25.2B MoE built on Gemma 4 that generates text in parallel for up to 4x faster output.
Cohere North Mini Code 1.0 is an open-weight 30B MoE coding model (3B active, 256K context, Apache 2.0) built for agentic software engineering. Specs, benchmarks, access.
Qwen3.7-Max is the closed flagship with higher vendor benchmarks and a 1M context; Kimi K2.7 Code is the open-weights, cheaper agentic specialist. We compare access, benchmarks, pricing, local feasibility, and which to use for autonomous coding.
Ornith 1.0 is a free, MIT-licensed, self-hostable coding model. Opus 4.8 is the closed frontier flagship. A benchmark-grounded, harness-honest comparison of where each wins on agentic coding in 2026.
DeepSeek V4 and Qwen 3.7 post near-identical coding benchmarks, but only one is actually open. A specifics-first comparison of architecture, local-run feasibility, API pricing, and license for developers choosing a coding model in 2026.
A realistic, no-hype guide to running Qwen 3.6 27B locally as a Claude Code alternative: benchmarks vs Opus, the hardware and quant you actually need, how to wire it in, where it holds up, and what a hybrid setup actually saves you.
Two new MIT open-weights coding models shipped a day apart in June 2026. We compare architecture, coding benchmarks, local hardware, and API pricing for Ornith 1.0 vs GLM 5.2 — with an honest, no-hype verdict on which to pick.
Zhipu Z.ai shipped GLM 5.2 today on every GLM Coding Plan tier with a usable 1M-token context window. Standalone API, the Z.ai chatbot, and the MIT open weights are arriving next week. No benchmarks yet — here's what's confirmed, what's not, and how it fits next to GLM-5.1.
JetBrains released Mellum2, a 12B Mixture-of-Experts model that activates just 2.5B parameters per token and ships under Apache 2.0. Here's what it is, where it fits in an AI stack, and how to put it to work.
A practical, size-tier-by-tier comparison of Google's Gemma 4 and Alibaba's Qwen 3.5 — benchmarks, coding, reasoning, multilingual, and how to run each locally in 2026.
Qwen 3.7 weights are not on Hugging Face yet (May 20, 2026). Here are the honest ways to use it today, and exactly what to run locally instead.
Qwen 3.6 is shipping with open weights today. Qwen 3.7-Max was announced May 20 with previews live but no weights yet. A grounded side-by-side.
MiniMax M3 is not released as of May 2026. Here's what's actually shipping (M2.7), where the 'M3.0 released' claim came from, and how to verify it.
Is Qwen 3.7 released? As of May 2026 it isn't — no weights, API, or benchmarks. Here's what's real, what's only rumored, and what to run today.
SubQ claims to be the first fully subquadratic LLM with a 12M-token context window. Here's what's verified, what isn't, and why the architecture matters.
ZAYA1-8B is an Apache-2.0 MoE reasoning model with 760M active params, pretrained 100% on AMD MI300X GPUs with zero NVIDIA in the loop.