GPT-5.6 Sol, Terra & Luna: Tiers, Pricing, Benchmarks
OpenAI's GPT-5.6 family — Sol, Terra, Luna — tiers, current pricing after the July 30 cuts, Ultrafast and Cyber, benchmarks, and how to choose.
A collection of 116 posts
OpenAI's GPT-5.6 family — Sol, Terra, Luna — tiers, current pricing after the July 30 cuts, Ultrafast and Cyber, benchmarks, and how to choose.
The 2026 memory crunch reshuffled the math on local AI. Here are the cheapest viable paths to run a local LLM right now — used GPUs, used Apple Silicon, CPU+RAM for MoE models, and cloud rental — ranked by dollars, with an honest beginner verdict.
Claude Fable 5 is back online as of July 1, 2026, after the U.S. lifted its export-control order. Here's what changed, how Anthropic brought it back, and how to access it.
Yes, gpt-3.5-turbo is still available — until October 23, 2026 (instruct: Sept 28). Exact shutdown dates, and why OpenAI points migrations at GPT-5.6 Terra.
OpenAI's GPT-5.5 Cyber ships gated under the Daybreak program. What 'trusted access' means, the CyberGym-vs-Mythos-5 benchmark claim (and its caveats), and what defenders and developers should take from it.
GLM-5.2 leads MiniMax M3 by 4.5 points on Artificial Analysis's Intelligence Index v4.3.2 — but only at its Max reasoning tier. M3 wins on price and owns image and video input. GLM-5.3-Flash now beats both on capability and cost.
A practical walkthrough for self-hosting GLM-5.2 (744B MoE, 40B active) on llama.cpp. Quant tables, four hardware paths, exact install commands, verification, and a fallback to the Z.ai cloud API if your rig falls short.
Four credible 128GB-class boxes, four very different price points. We synthesise what practitioners with the hardware on their desks are actually reporting.
VibeThinker-3B is WeiboAI's MIT-licensed 3B reasoning model built on Qwen2.5-Coder-3B. We unpack the viral 'Opus 4.5 performance' claim with the actual HF benchmarks.
Z.ai's GLM-5.2: 744B params (40B active), 1M-token context, MIT-licensed weights — still the newest GLM you can self-host while GLM-5.3's weights are pending. Architecture, benchmarks, pricing, and a 3-path local-inference playbook.
Two open-weights heavyweights from China go head-to-head for the agentic-coding throne. K2.7 leads on MCP tool-use depth; V4 leads on raw per-token economics and proven independent benchmarks. We break down cost, agentic strength, self-host paths, and pick a winner per workload.
Moonshot's open-weights Kimi K2.7 Code goes head-to-head with Anthropic's Claude Opus 4.8. Architecture, benchmarks (and where they don't exist yet), per-task cost, agentic strength, self-host paths, and a clean per-workload verdict.
Moonshot's Kimi K2.7 Code and Z.ai's freshly-released GLM 5.2 are both Chinese open-weights coding flagships, both shipped in June 2026, and they trade on opposite axes. K2.7 leads on MCP tool use and pricing; GLM 5.2 leads on 1M context. We pick per workload.
Two open-weights heavyweights from China go head-to-head for the agentic-coding throne. GLM 5.2 leads on context window; DeepSeek V4 leads on token economics. We break down cost, agentic strength, self-host paths, and pick a winner per workload.
Claude Opus 4.8 leads GLM 5.2 on Artificial Analysis's Intelligence Index v4.3.2 (41.79 vs 33.71), and both carry 1M context. GLM 5.2's real edge is MIT weights — and GLM-5.3-Flash now matches Opus 4.8 at a sixteenth of the cost per task.
GPT-5.5 leads GLM 5.2 on Artificial Analysis's Intelligence Index v4.3.2 (38.36 vs 33.71), and both carry roughly 1M context. The per-task cost gap is 1.8x, not the 6.8x the per-token rates imply.
Zhipu Z.ai shipped GLM 5.2 today on every GLM Coding Plan tier with a usable 1M-token context window. Standalone API, the Z.ai chatbot, and the MIT open weights are arriving next week. No benchmarks yet — here's what's confirmed, what's not, and how it fits next to GLM-5.1.
How Moonshot's open-weight Kimi K2.7 Code stacks up against Claude Opus 4.8, GPT-5.5, and DeepSeek V4 for agentic coding — on price, context, and the benchmarks that exist. K2.7's scores are Moonshot-reported only, so the verdict is subject to change once independent results land.
Moonshot AI's Kimi K2.7 Code — a 1T-parameter open-weight coding model with a 256K context, ~30% fewer thinking tokens than K2.6, and strong MCP tool-use. Benchmarks, pricing, API, and local-deployment guide.
Anthropic's first publicly available Mythos-class model, released June 9, 2026. Third-party benchmarks, pricing, context window, availability, the safety reroute to Opus 4.8, and how it compares to GPT-5.5 and Gemini 3.5.
A practical 2026 comparison of Kimi K2.6, GPT-5.5, and Claude Opus 4.8 on coding benchmarks, reasoning, pricing, and self-host economics — plus which to pick by use case.
May 2026 state of the AI benchmark leaderboard: SWE-bench Verified + Pro, GAIA, Terminal-Bench 2.0, GDPval, MCP Atlas, USAMO, GPQA, HLE. Who leads, what's the gap, what each score actually means.
Anthropic launched Claude Opus 4.8 on May 28, 2026: SWE-bench Pro 69.2%, GDPval Elo 1890 (+121 over GPT-5.5), Fast mode 3x cheaper than 4.7, dynamic workflows for hundreds of parallel subagents. Pricing unchanged at $5/$25 per 1M. Full launch breakdown.
Two weeks after Qwen 3.7 Max, Alibaba shipped WebWorld: an Apache 2.0 web world model series that simulates browsers for agent training. Sizes, benchmarks, code, gotchas.
xAI launched Grok Imagine Agent Mode on May 1, 2026 — an infinite-canvas creative agent that plans, generates, edits, and stitches 6-second video clips into longer films. Features, four templates, vs Sora and Veo, pricing, and API examples.