Tag

Qwen

A collection of 32 posts

AI Models

Qwen Image 2.1: Run It Locally (ComfyUI, GGUF, Diffusers) 2026

Run Qwen-Image-2.1 locally on 8-24GB GPUs: official ComfyUI files, GGUF quants for low VRAM, and the Diffusers pipeline. Includes a VRAM table, transparent PNG and multi-reference editing, the vendor benchmark, and the non-commercial license caveat.

· 14 min read
Qwen

Qwen 3.8 Flash: 1M Context at $0.15/M (2026 Guide)

Alibaba's cheap multimodal tier: $0.15/$0.47 per 1M tokens, 1M context, and open weights under the Qwen Community License. Specs, benchmarks and how it compares to GLM-5.3-Flash and DeepSeek V4-Flash.

· 11 min read
Qwen

Qwen3.8-27B as a Local Claude Code Replacement (2026)

Qwen3.8-27B posts benchmark wins over Claude Opus 4.6 Max and ships a first-party Claude Code launcher via Ollama. We grade the claims, fix the overthinking latency problem, and redraw the hybrid local-vs-cloud line.

· 11 min read
Qwen

Qwen 3.8 vs Qwen 3.6: Same Architecture, +12 Points (2026)

The config.json diff between Qwen 3.8-27B and Qwen 3.6-27B is empty — same layers, same hidden size, same vocab. Every gain came from post-training. Here's what actually changed, what the upgrade costs you, and who should stay on 3.6.

· 7 min read
Qwen

Qwen 3.7 vs Kimi K2.7: Best Open Agentic Coder in 2026?

Qwen3.7-Max is the closed flagship with higher vendor benchmarks and a 1M context; Kimi K2.7 Code is the open-weights, cheaper agentic specialist. We compare access, benchmarks, pricing, local feasibility, and which to use for autonomous coding.

· 13 min read
Qwen

Qwen 3.6 27B as a Local Claude Code Replacement

A realistic, no-hype guide to running Qwen 3.6 27B locally as a Claude Code alternative: benchmarks vs Opus, the hardware and quant you actually need, how to wire it in, where it holds up, and what a hybrid setup actually saves you.

· 16 min read
AI

Qwen 3.7 Max: Alibaba's May 2026 Flagship Guide

Alibaba's Qwen 3.7 Max launched May 20, 2026 with a 1M-token context, native extended-thinking mode, and benchmark wins on SWE-Pro and Terminal-Bench. Here's how it compares to Claude Opus 4.7, GPT-5.5, Gemini 3.5 Flash and DeepSeek V4, what it costs on DashScope, and when to pick it.

· 11 min read
Qwen3-VL-4B Instruct vs Qwen3-VL-4B Thinking: Complete 2026 Guide
Qwen

Qwen3-VL-4B Instruct vs Qwen3-VL-4B Thinking: Complete 2026 Guide

Quick answer. Qwen3-VL-4B Instruct and Thinking share a 4.44B dense transformer (256K context, 1M expandable). Pick Instruct for fast multimodal chat at 55-75 tok/s FP8 on a 12 GB GPU; pick Thinking for math, multi-step reasoning, and long video where 94.2% DocVQA matters more than speed. Last

· 20 min read
Qwen3-VL-8B Instruct vs Qwen3-VL-8B Thinking: 2026 Guide
AI

Qwen3-VL-8B Instruct vs Qwen3-VL-8B Thinking: 2026 Guide

Quick answer. Qwen3-VL-8B Instruct and Thinking share the same 9B Apache 2.0 backbone and differ only in post-training. Pick Instruct for high-volume OCR, chatbots, and production pipelines at roughly 45-60 tok/s on a 4090. Pick Thinking for STEM, medical, legal, or mockup-to-code tasks where the 2-4 point benchmark

· 16 min read
Qwen3-VL-30B-A3B-Thinking: Complete 2026 Deployment Guide
AI

Qwen3-VL-30B-A3B-Thinking: Complete 2026 Deployment Guide

Master Qwen3-VL-30B-A3B-Thinking deployment with our comprehensive 2025 guide. Learn installation, optimization, troubleshooting, and real-world applications for this powerful 30B parameter vision-language AI model with thinking capabilities.

· 17 min read