MiMo V2.6 Flash & 9B Distill: Benchmarks, Run Locally (2026)
Xiaomi put the MiMo-V2.6 family on Hugging Face on 21 September 2026 (some coverage dates the launch 22 September). Everything is MIT-licensed. The launch had two open checkpoints: MiMo-V2.6-Pro-RL and MiMo-V2.6-Flash-RL. Next to them was a smaller release that most local-LLM users will care about more: MiMo-V2.6-Distill-Qwen-9B, which is Alibaba's Qwen3.5-9B fine-tuned on data generated by MiMo. On 27 September Xiaomi followed up with -Flash-MOPD and -Pro-MOPD checkpoints. These fix a failure mode where the model kept sending the same tool call over and over.
This guide covers the two models people keep searching for, Flash and the 9B distill. You get the specs from the official model cards, the vendor benchmarks next to Claude Opus 5 and the Qwen3.5-9B base, the independent Artificial Analysis numbers, and working commands for running the 9B locally in Ollama, LM Studio and llama.cpp. There is also a VRAM table built from the actual GGUF file sizes. If you want the flagship instead, read our MiMo-V2.6-Pro guide.
What is MiMo-V2.6-Flash?
MiMo-V2.6-Flash is the "efficiency-balanced checkpoint" of the V2.6 series, in Xiaomi's words. It is a sparse Mixture-of-Experts model with 309B total parameters and 15B active per token. It has 256 routed experts, 8 of which are active at a time, and no shared experts. It accepts text, image, video and audio in one model, and the context window is 1M tokens. According to the model card, the whole series was trained with one mixed reinforcement-learning run across coding, general agent, visual and cybersecurity tasks, rather than separate runs per domain. Xiaomi calls this "You Only RL Once".
The backbone is a hybrid. Of its 48 layers, 39 use sliding-window attention with a 128-token window and 9 use global attention. That split is what makes a 1M context affordable at inference time. The model also ships a 5-layer multi-token-prediction (MTP) speculative decoder that predicts 7 tokens ahead per forward pass.
MiMo-V2.6-Flash specs
| Spec | MiMo-V2.6-Flash-RL |
|---|---|
| Release | 21 Sep 2026 (Hugging Face) |
| License | MIT |
| Architecture | Sparse MoE, 309B total / 15B active |
| Experts | 256 routed, 8 active, no shared experts |
| Layers | 48 (39 sliding-window + 9 global attention) |
| Context length | 1M tokens |
| Input modalities | Text, image, video, audio |
| Vision encoder | 681M-param MiMo ViT (28 layers) |
| Audio encoders | 308M AudioTokenizer + 127M audio patch encoder |
| Speculative decoding | 5-layer MTP drafter, 7 tokens per pass |
| Recommended sampling | temperature 1.0, top_p 0.95 |
| API price (OpenRouter) | $0.14 input / $0.28 output per 1M tokens |
How good is MiMo-V2.6-Flash? Benchmarks vs Pro and Claude Opus 5
The table below uses Xiaomi's own numbers from the Flash-RL model card, so treat them as vendor-reported. According to those numbers, Flash stays within about 1-4 points of the much larger Pro on most agent benchmarks and within about 1.5-7 points of Claude Opus 5 on most of them. On AutomationBench it beats Opus 5. It falls well behind on the hardest long-horizon and exploit tests.
| Benchmark | MiMo-V2.6 Flash | MiMo-V2.6 Pro | Claude Opus 5 | Source |
|---|---|---|---|---|
| DeepSWE v1.1 | 67.9 | 71.9 | 74.0 | Xiaomi (vendor) |
| Terminal Bench 2.1 | 87.6 | 89.9 | 89.1 | Xiaomi (vendor) |
| Terminal Bench 4.0 | 28.8 | 34.9 | 49.0 | Xiaomi (vendor) |
| OSWorld-Verified | 80.8 | 82.0 | 83.4 | Xiaomi (vendor) |
| Toolathlon-Verified | 73.6 | 76.9 | 80.6 | Xiaomi (vendor) |
| AutomationBench v1.0.6 | 52.3 | 53.1 | 50.3 | Xiaomi (vendor) |
| JobBench | 61.2 | 62.0 | 65.7 | Xiaomi (vendor) |
| ProgramBench | 26.0 | 26.5 | 37.0 | Xiaomi (vendor) |
| ExploitBench | 25.3 | 47.9 | 70.0 | Xiaomi (vendor) |
| Artificial Analysis Intelligence Index | 38 (#8 of 117) | 46 (#1 of 117) | n/a | Artificial Analysis (independent) |
Artificial Analysis gives an independent check. Pro scores 46 on its Intelligence Index, first of the 117 open-weight models it compares. Flash scores 38, which is eighth. Artificial Analysis also measured Flash at 55.4 output tokens per second, which is below average for its class, and about $0.06 per task on its evaluation suite. So the independent gap between Flash and Pro is larger than the vendor agent tables suggest. On agentic tasks the two are close. On broad reasoning, Pro is clearly ahead.
How much does MiMo-V2.6-Flash cost?
On OpenRouter, xiaomi/mimo-v2.6-flash costs $0.14 per million input tokens and $0.28 per million output tokens with a context of about 1.05M tokens. That is about a third of Pro ($0.435 / $0.87), and the same price OpenRouter lists for MiMo-V2.5. Xiaomi also sells access directly through the MiMo API Platform.
Here is a minimal OpenAI-compatible call through OpenRouter:
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "xiaomi/mimo-v2.6-flash",
"messages": [{"role": "user", "content": "Write a bash one-liner that finds the 10 largest files in a repo."}]
}'Can you run MiMo-V2.6-Flash locally?
Not on normal hardware. Xiaomi's reference SGLang command for Flash uses tensor parallel 8 (--tp 8), which means one 8-GPU node. Its vLLM example uses 4-way tensor parallel. The smallest GGUFs show the problem clearly. In ggml-org/MiMo-V2.6-Flash-RL-GGUF, the Q2_K build is about 126 GB and the MXFP4 build is about 167 GB, and that is before you add the vision/audio projector or KV cache. A workstation with 192 GB or more of unified memory or multiple 96 GB cards can load it. For everyone else, Flash is an API model.
If you do self-host on a GPU node, this is the vLLM command from the model card:
vllm serve XiaomiMiMo/MiMo-V2.6-Flash-RL \
--tensor-parallel-size 4 \
--trust-remote-code \
--gpu-memory-utilization 0.95 \
--max-model-len auto \
--reasoning-parser mimo \
--tool-call-parser mimo \
--enable-auto-tool-choice \
--generation-config vllmOur self-hosting LLMs guide covers serving stacks, quantization trade-offs and GPU sizing in more depth.
What are the Flash-MOPD and Pro-MOPD checkpoints?
After launch, users reported that V2.6 sometimes kept issuing the same or nearly the same tool call. The agent looked busy and made no progress. Xiaomi's technical blog measured the repetition rate before the fix at 0.07% to 1.02% for Flash-RL and 0.05% to 0.54% for Pro-RL, depending on the agent harness. The worst rates were in OpenCode. The rates sound small, but one loop can burn a long agent run's time and context.
The fix was a short MOPD (multi-teacher on-policy distillation) pass. Xiaomi says it cost about $90,000, around 4% of the estimated $2.31 million for the alternative of fixing it with another mixed RL run. After it, Xiaomi reports that repetition "dropped substantially across context lengths and agent harnesses", with many cells in its heatmaps at or below 0.001%. The fixed weights were published on 27 September as MiMo-V2.6-Flash-MOPD and Pro-MOPD, with the same architecture and specs. If you are self-hosting Flash for agent work, use the MOPD checkpoint. ggml-org also publishes MiMo-V2.6-Flash-MOPD-GGUF.
What is MiMo-V2.6-Distill-Qwen-9B?
MiMo-V2.6-Distill-Qwen-9B is a supervised fine-tune of Qwen3.5-9B, trained on 77.4B tokens of MiMo-generated data (27.2B of them loss-bearing). By total tokens the mix was code 29.9%, general agent tasks 28.5%, visual tasks 27.4% and cybersecurity 14.2%. It inherits the Qwen3.5 architecture: 32 layers in which every fourth layer uses full attention and the rest use linear attention, a 262,144-token configured context, and image input through a vision projector. It has 9.41B parameters in BF16.
Xiaomi describes it as an SFT checkpoint released "as a starting point for open research in agentic reinforcement learning". It is not a tuned product model, and eesel's review calls it "a research starting point rather than production". It is still the only way to get MiMo-flavoured agent behaviour on a laptop, which is why its GGUFs have outpaced every official MiMo repo in downloads. bartowski's GGUF repo had about 205,000 downloads when we checked, compared with about 12,000 for Xiaomi's own safetensors.
MiMo-V2.6-Distill-Qwen-9B vs Qwen3.5-9B benchmarks
| Benchmark (metric) | Qwen3.5-9B (base) | MiMo-V2.6-Distill-Qwen-9B | Change |
|---|---|---|---|
| SWE Verified (avg@3) | 60.0 | 61.1 | +1.1 |
| SWE Pro (avg@3) | 32.0 | 44.6 | +12.6 |
| Terminal Bench 2.1 (avg@1) | 27.0 | 37.1 | +10.1 |
| Toolathlon-Verified (avg@1) | 25.9 | 35.2 | +9.3 |
| AutomationBench v1.0.6 (avg@1) | 5.0 | 30.3 | +25.3 |
| OfficeQA (avg@1) | 9.0 | 19.5 | +10.5 |
| JobBench (avg@1) | 2.6 | 18.3 | +15.7 |
All of these numbers come from Xiaomi's technical report as reproduced on the model card, so they are vendor-reported. We left out the "mini" internal MiMo benchmarks because nobody outside Xiaomi can reproduce them. The pattern is clear. On SWE Verified, a bug-fixing benchmark where Qwen3.5-9B was already strong, the distill is flat. On longer, multi-step agent work (SWE Pro, Terminal Bench, tool use, automation), it gains 9 to 25 points. That matches what the training data targeted.
For scale, the 9B's 37.1 on Terminal Bench 2.1 compares with 87.6 for Flash. It is a much smaller model and does not come close. Its value is that it runs offline on hardware you already own.
How much VRAM does MiMo-V2.6 9B need?
The file sizes below come from the Hugging Face file listing for bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF. The memory columns are our estimates: file size plus KV cache plus about 0.7 GB of runtime and compute buffers. The KV cache is small because only 8 of the 32 layers use full attention (4 KV heads, head dim 256). That works out to about 32 KB per token at FP16: roughly 0.27 GB at 8K context, 1.1 GB at 32K and 4.3 GB at 128K. For image input, add the mmproj file (0.92 GB for bartowski's, 0.62 GB for ggml-org's Q8_0).
| Quant | File size | Est. memory @ 8K ctx | Est. memory @ 32K ctx | Fits on |
|---|---|---|---|---|
| IQ2_M | 3.54 GB | ~4.5 GB | ~5.4 GB | 6 GB GPU (quality drops noticeably) |
| Q3_K_M | 4.48 GB | ~5.5 GB | ~6.3 GB | 6-8 GB GPU |
| IQ4_XS | 5.23 GB | ~6.2 GB | ~7.0 GB | 8 GB GPU |
| Q4_K_M (recommended) | 5.84 GB | ~6.8 GB | ~7.6 GB | 8 GB GPU, 16 GB Mac |
| Q5_K_M | 6.88 GB | ~7.9 GB | ~8.7 GB | 10-12 GB GPU |
| Q6_K | 7.79 GB | ~8.8 GB | ~9.6 GB | 12 GB GPU |
| Q8_0 | 9.55 GB | ~10.5 GB | ~11.4 GB | 12-16 GB GPU, 16 GB Mac (tight) |
| BF16 | 17.92 GB | ~18.9 GB | ~19.7 GB | 24 GB GPU |
Q4_K_M is bartowski's default recommendation and a good starting point. If you are going to rely on the model for agentic coding, where small errors compound over many steps, use Q6_K or Q8_0 if you have the memory. Speeds reported so far are modest. Atomic Chat collected about 6 tokens/s generation on a 16 GB M4 MacBook Pro in one test and 13.7-15.3 tokens/s on an 18 GB M3 Pro in another. The two tests used different quants and contexts, so do not compare them directly.
How to run MiMo-V2.6 9B locally
Ollama has no official library tag for MiMo-V2.6. There is no ollama run mimo. There are community uploads under user namespaces (for example maternion/mimo-v2.6), but the cleaner route is to pull the GGUF straight from Hugging Face, which Ollama supports natively. Whichever runtime you use, it needs Qwen3.5 architecture support. Ollama already ships qwen3.5 in its library. For llama.cpp, bartowski quantized with release b10964 and recommends that release or newer.
Ollama
# Q4_K_M from bartowski (5.84 GB)
ollama run hf.co/bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q4_K_M
# Higher quality if you have ~10 GB free
ollama run hf.co/bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q6_K
# ggml-org's official conversion (Q8_0 only)
ollama run hf.co/ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q8_0Ollama's default context is small. For coding agents, raise it with /set parameter num_ctx 32768 in the REPL or PARAMETER num_ctx 32768 in a Modelfile. As the VRAM table shows, 32K costs only about 1 GB extra on this model.
LM Studio
In LM Studio, open the Discover tab and search bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF. Pick Q4_K_M or higher, and load it with an 8K-32K context. bartowski lists LM Studio as compatible. LM Studio has no staff-picked MiMo listing yet, so use the bartowski repo. Download the mmproj file from the same repo too if you want image input. On Apple Silicon, mlx-community/MiMo-V2.6-Distill-Qwen-9B-OptiQ-4bit is an MLX alternative.
llama.cpp
# One-line install (from bartowski's model card), then serve
curl -LsSf https://llama.app/install.sh | sh
llama-server -hf bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q4_K_M
# Or download a specific quant first
hf download bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF \
--include "MiMo-V2.6-Distill-Qwen-9B-Q4_K_M.gguf" --local-dir ./llama-server serves a chat UI and an OpenAI-compatible API at http://localhost:8080. The -hf flag fetches the mmproj automatically, so image input works without extra steps.
Settings that matter
- Sampling: the model's
generation_config.jsonsets temperature 0.6, top_k 20, top_p 0.95. Flash's card recommends temperature 1.0, but that setting is for Flash, not the 9B. - Thinking mode: this is a reasoning model. Xiaomi's quickstart toggles thinking with
chat_template_kwargs: {"enable_thinking": true}. Set it tofalsefor faster, shorter replies. - Tool calling: ggml-org's card notes that a chat-template patch is needed until the upstream template is fixed. Atomic Chat reports that a MiMo tool-parser fix landed in llama.cpp b11102. If tool calls come back malformed, update llama.cpp before you debug your agent.
- Full precision serving: Xiaomi's official route for the unquantized model is SGLang:
sglang serve --model-path XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B --reasoning-parser mimo.
MiMo-V2.6 9B vs Qwen 3.8: which small model should you run?
If you already run a Qwen model locally (see our guide to running Qwen 3.8 locally), the question is whether the distill beats your current default. Xiaomi only compares it with its own base, Qwen3.5-9B. It does not compare it with newer Qwen releases, so there is no apples-to-apples vendor table. Our practical advice: if your GPU fits a larger model, a mid-size model such as Qwen 3.8 27B will generally beat any 9B on hard coding. The MiMo distill makes sense when you are limited to 8-12 GB and want better multi-step tool use than a stock 9B gives you. For the broader field, see our open-source LLM landscape.
Who should use MiMo-V2.6-Flash or the 9B?
- Teams running agent pipelines on a budget: use Flash through the API. At $0.14/$0.28 it is one of the cheapest models that scores in the high 80s on Terminal Bench 2.1 (vendor-reported). Test the MOPD checkpoint if you self-host.
- Teams needing the best open-weights quality: use Pro. The independent Intelligence Index gap (46 vs 38) is real.
- Developers who want a private, offline coding helper on an 8-16 GB machine: use the 9B distill at Q4_K_M or Q6_K. Expect a capable assistant, not an autonomous engineer.
- Researchers working on agentic RL: the 9B SFT checkpoint is explicitly meant to be a starting point for your own RL runs.
- Anyone who needs audio or video understanding locally: neither option works. The 9B takes text and images only, and omnimodal Flash needs server hardware.
If you used earlier MiMo releases, our MiMo-V2.5 coding model write-up shows how far the line has moved. V2.5 Pro scored 19.0 on DeepSWE v1.1, and V2.6 Flash scores 67.9.
FAQ
Is MiMo-V2.6-Flash open source?
Yes. The weights are on Hugging Face under the MIT license, which allows commercial use and fine-tuning. The same applies to Pro, the MOPD checkpoints and the Distill-Qwen-9B.
Is there an official Ollama model for MiMo-V2.6?
No. As of 30 September 2026 there is no MiMo entry in the Ollama library, only community uploads. Run the GGUF directly with ollama run hf.co/bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q4_K_M.
How much VRAM does MiMo-V2.6-Distill-Qwen-9B need?
About 7 GB at Q4_K_M with an 8K context, and about 8 GB at 32K. Q8_0 needs about 10.5-11.5 GB, and BF16 needs about 19-20 GB. The KV cache is unusually cheap because only a quarter of the layers use full attention.
Can I run MiMo-V2.6-Flash on a single GPU?
Not in any practical way. Even the Q2_K GGUF is about 126 GB. Xiaomi's reference deployment uses 8-way tensor parallelism on one node. Use the API or a multi-GPU server.
What is the difference between Flash-RL and Flash-MOPD?
MOPD is a follow-up checkpoint released on 27 September 2026. It adds a multi-teacher on-policy distillation pass that sharply reduces repeated tool calls in agent loops. The architecture and specs are the same, and it is the better choice for agent work.
Does the 9B support images and long context?
It supports image input through an mmproj file, which llama.cpp downloads automatically with -hf. The configured context is 262,144 tokens, but on consumer hardware 8K-32K is the realistic range.
How does MiMo-V2.6-Flash compare with Claude Opus 5?
On Xiaomi's own benchmarks, Flash is 1.5-7 points behind Opus 5 on most agent tests, slightly ahead on AutomationBench (52.3 vs 50.3), and far behind on Terminal Bench 4.0, ProgramBench and exploit tasks. It costs a small fraction of the price.
Sources
- XiaomiMiMo/MiMo-V2.6-Flash-RL model card (Hugging Face)
- XiaomiMiMo/MiMo-V2.6-Flash-MOPD model card (Hugging Face)
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B model card and config (Hugging Face)
- bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUF (file sizes, run instructions)
- ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF
- ggml-org/MiMo-V2.6-Flash-RL-GGUF (Flash quant sizes)
- Xiaomi MiMo blog: Diagnosing and mitigating tool-call repetition in MiMo-V2.6
- OpenRouter: MiMo-V2.6-Flash pricing
- Artificial Analysis: MiMo-V2.6-Flash
- Artificial Analysis: MiMo-V2.6-Pro
- eesel AI: Xiaomi MiMo V2.6 review
- Atomic Chat: How to run MiMo-V2.6 9B locally
If your team is building products on open models like MiMo, Codersera can help you hire vetted remote developers who have shipped LLM-backed features.