Tag

LLM

A collection of 116 posts

AI

GLM-5.2 vs MiniMax M3: Open-Weights Coding (2026)

GLM-5.2 leads MiniMax M3 by 4.5 points on Artificial Analysis's Intelligence Index v4.3.2 — but only at its Max reasoning tier. M3 wins on price and owns image and video input. GLM-5.3-Flash now beats both on capability and cost.

· 11 min read
AI

How to Run GLM-5.2 Locally — and When to Pick 5.3-Flash

A practical walkthrough for self-hosting GLM-5.2 (744B MoE, 40B active) on llama.cpp. Quant tables, four hardware paths, exact install commands, verification, and a fallback to the Z.ai cloud API if your rig falls short.

· 14 min read
AI

VibeThinker-3B: The Complete Guide (2026)

VibeThinker-3B is WeiboAI's MIT-licensed 3B reasoning model built on Qwen2.5-Coder-3B. We unpack the viral 'Opus 4.5 performance' claim with the actual HF benchmarks.

· 9 min read
AI

GLM-5.2: 744B MoE, 1M Context, MIT-Licensed (2026)

Z.ai's GLM-5.2: 744B params (40B active), 1M-token context, MIT-licensed weights — still the newest GLM you can self-host while GLM-5.3's weights are pending. Architecture, benchmarks, pricing, and a 3-path local-inference playbook.

· 18 min read
AI

Kimi K2.7 vs DeepSeek V4: The Open-Weights Coding Showdown (2026)

Two open-weights heavyweights from China go head-to-head for the agentic-coding throne. K2.7 leads on MCP tool-use depth; V4 leads on raw per-token economics and proven independent benchmarks. We break down cost, agentic strength, self-host paths, and pick a winner per workload.

· 6 min read
AI

Kimi K2.7 vs Claude Opus 4.8: Coding Compared (2026)

Moonshot's open-weights Kimi K2.7 Code goes head-to-head with Anthropic's Claude Opus 4.8. Architecture, benchmarks (and where they don't exist yet), per-task cost, agentic strength, self-host paths, and a clean per-workload verdict.

· 9 min read
AI

Kimi K2.7 vs GLM 5.2: Open-Weights Coding Compared

Moonshot's Kimi K2.7 Code and Z.ai's freshly-released GLM 5.2 are both Chinese open-weights coding flagships, both shipped in June 2026, and they trade on opposite axes. K2.7 leads on MCP tool use and pricing; GLM 5.2 leads on 1M context. We pick per workload.

· 14 min read
AI

GLM 5.2 vs DeepSeek V4: The Open-Weights Coding Showdown (2026)

Two open-weights heavyweights from China go head-to-head for the agentic-coding throne. GLM 5.2 leads on context window; DeepSeek V4 leads on token economics. We break down cost, agentic strength, self-host paths, and pick a winner per workload.

· 15 min read
AI

GLM 5.2 vs Claude Opus 4.8 for Coding (2026)

Claude Opus 4.8 leads GLM 5.2 on Artificial Analysis's Intelligence Index v4.3.2 (41.79 vs 33.71), and both carry 1M context. GLM 5.2's real edge is MIT weights — and GLM-5.3-Flash now matches Opus 4.8 at a sixteenth of the cost per task.

· 15 min read
AI

GLM 5.2 vs GPT-5.5 for Coding: Open vs Closed (2026)

GPT-5.5 leads GLM 5.2 on Artificial Analysis's Intelligence Index v4.3.2 (38.36 vs 33.71), and both carry roughly 1M context. The per-task cost gap is 1.8x, not the 6.8x the per-token rates imply.

· 14 min read
Kimi

Kimi K2.7 vs GPT-5.5 vs Claude Opus 4.8: Coding & Agentic Comparison (2026)

How Moonshot's open-weight Kimi K2.7 Code stacks up against Claude Opus 4.8, GPT-5.5, and DeepSeek V4 for agentic coding — on price, context, and the benchmarks that exist. K2.7's scores are Moonshot-reported only, so the verdict is subject to change once independent results land.

· 9 min read
Claude

Claude Fable 5: Release Date, Features, Pricing (2026)

Anthropic's first publicly available Mythos-class model, released June 9, 2026. Third-party benchmarks, pricing, context window, availability, the safety reroute to Opus 4.8, and how it compares to GPT-5.5 and Gemini 3.5.

· 11 min read
AI Models

Kimi K2.6 vs GPT-5.5 vs Claude Opus 4.8 (2026)

A practical 2026 comparison of Kimi K2.6, GPT-5.5, and Claude Opus 4.8 on coding benchmarks, reasoning, pricing, and self-host economics — plus which to pick by use case.

· 8 min read
AI

Claude Opus 4.8 Launch Guide: Benchmarks & Pricing 2026

Anthropic launched Claude Opus 4.8 on May 28, 2026: SWE-bench Pro 69.2%, GDPval Elo 1890 (+121 over GPT-5.5), Fast mode 3x cheaper than 4.7, dynamic workflows for hundreds of parallel subagents. Pricing unchanged at $5/$25 per 1M. Full launch breakdown.

· 12 min read