Run Baidu Unlimited-OCR Locally: Transformers, vLLM & SGLang (2026 Guide)
Baidu's Unlimited-OCR parses entire multi-page PDFs in a single forward pass. Here's how to run the 3.3B open-weights model locally with Transformers, vLLM, or SGLang.
A collection of 32 posts
Baidu's Unlimited-OCR parses entire multi-page PDFs in a single forward pass. Here's how to run the 3.3B open-weights model locally with Transformers, vLLM, or SGLang.
DSpark is DeepSeek's open-source speculative-decoding module that makes V4-Pro and V4-Flash 51–400% faster — and it works on Qwen3 and Gemma 4 too. Here's how it works and how to use it.
OpenCode just passed Claude Code on GitHub stars. We compare the open-source, model-agnostic terminal agent against Anthropic's polished CLI on setup, models, MCP, pricing, and who should pick which.
A coding head-to-head: GLM-5.2's leaderboard-topping text coding vs MiniMax M3's native multimodality, lower price, and MSA long-context speed — with specs, benchmarks, and a clear verdict.
VibeThinker-3B is WeiboAI's MIT-licensed 3B reasoning model built on Qwen2.5-Coder-3B. We unpack the viral 'Opus 4.5 performance' claim with the actual HF benchmarks.
Z.ai's GLM-5.2 is the leading open-weights LLM on the Artificial Analysis Intelligence Index v4.1. 744B params (40B active), 1M-token context, MIT-licensed weights. Architecture, benchmarks, pricing, and a 3-path local-inference playbook.
Two open-weights heavyweights from China go head-to-head for the agentic-coding throne. K2.7 leads on MCP tool-use depth; V4 leads on raw per-token economics and proven independent benchmarks. We break down cost, agentic strength, self-host paths, and pick a winner per workload.
Moonshot's Kimi K2.7 Code and Z.ai's freshly-released GLM 5.2 are both Chinese open-weights coding flagships, both shipped in June 2026, and they trade on opposite axes. K2.7 leads on MCP tool use and pricing; GLM 5.2 leads on 1M context. We pick per workload.
Two open-weights heavyweights from China go head-to-head for the agentic-coding throne. GLM 5.2 leads on context window; DeepSeek V4 leads on token economics. We break down cost, agentic strength, self-host paths, and pick a winner per workload.
Moonshot AI's Kimi K2.7 Code — a 1T-parameter open-weight coding model with a 256K context, ~30% fewer thinking tokens than K2.6, and strong MCP tool-use. Benchmarks, pricing, API, and local-deployment guide.
H Company's Holo3.1 family brings computer-use agents to local and on-device inference with quantized checkpoints and four model sizes. Here's what shipped and how to deploy it.
Two weeks after Qwen 3.7 Max, Alibaba shipped WebWorld: an Apache 2.0 web world model series that simulates browsers for agent training. Sizes, benchmarks, code, gotchas.
DeepSeek made its 75% V4-Pro discount permanent on May 22, 2026. Standing rates: $0.435/M input, $0.87/M output. Here is what changed, the new cost-per-quality math vs Claude Opus 4.7 and GPT-5.5, and the migration code.
Cohere released Command A+ on May 20, 2026: a 218B sparse Mixture-of-Experts model with 25B active parameters, Apache 2.0 licensed, that runs on as few as 2 H100 GPUs. Built for sovereign, on-prem enterprise agents with native citations.
Alibaba's Qwen 3.7 Max launched May 20, 2026 with a 1M-token context, native extended-thinking mode, and benchmark wins on SWE-Pro and Terminal-Bench. Here's how it compares to Claude Opus 4.7, GPT-5.5, Gemini 3.5 Flash and DeepSeek V4, what it costs on DashScope, and when to pick it.
Mistral Medium 3.5 is a 128B dense model with built-in reasoning, coding, and agentic capabilities. Le Chat Work Mode turns it into a multi-tool agent. Here's what's new, what it costs, and when to actually pick Mistral over Claude or GPT.
Void is a free, open-source, VS Code-based AI code editor and Cursor alternative. Here's what it does, how it works, and whether the paused project is worth using in 2026.
Kimi K2.6 and DeepSeek V4 Pro are the two best open-weights coding models in 2026. K2.6 wins long-horizon agents and swarms; DeepSeek V4 wins on raw price.
Five frontier-class open-weight LLMs shipped in 30 days. Real benchmarks, licenses, hosting costs, and a decision matrix for CTOs picking their 2026 stack.
DeepSeek V4 Flash is the under-covered story of the V4 release. 1M context, 47 on the AA Intelligence Index, $0.14 input / $0.28 output per million tokens, and it fits on a Mac Studio. Here is the full practical guide.
Claude Code's source is now public on GitHub. This guide covers what the OSS release actually means, every install method, project configuration, BYOK via LiteLLM, and power-user tips for MCP servers and GitHub Actions.
Learn how to install Void AI, the open-source Cursor alternative, and run it with local models via Ollama or LM Studio — with zero cloud dependencies.
A technical comparison of Void AI and Cursor covering privacy architecture, local model support, feature parity, pricing, and the development pause that changes Void's long-term outlook.
Void AI is an open-source, VS Code-based code editor that brings Cursor-style AI features — inline editing, agent mode, and autocomplete — without routing your code through a proprietary backend. Here's what it does and who should use it.
Sora's API is shutting down, Runway charges at scale, and Mochi 1 has quietly caught up on quality. Here's the practical comparison for developers building video pipelines.