Qwen Image 2.1: Run It Locally (ComfyUI, GGUF, Diffusers) 2026

Quick answer. Qwen-Image-2.1 is Alibaba's open-weight 7B image generation and editing model, released September 20, 2026. The easiest local route is ComfyUI with the Comfy-Org INT8 files on a 16-24GB GPU. On 8-12GB cards, use a Q4_K_M or Q5 GGUF through the leejet ComfyUI-GGUF node. The license is research-only, so commercial use needs a separate license from Qwen.

Qwen released Qwen-Image-2.1 on September 20, 2026, with weights on Hugging Face and ModelScope. It is one checkpoint that handles both text-to-image and image editing. It also outputs native transparent PNGs, accepts up to 10 reference images, and renders at 2K. Diffusers, ComfyUI, vLLM-Omni, SGLang and LightX2V all supported it on day one. Within ten days, the ComfyUI repackage on Hugging Face had passed 5 million downloads.

This guide covers three ways to run it on your own GPU: the official ComfyUI files, GGUF quantizations for 8-12GB cards, and the Python Diffusers pipeline. It also covers how much VRAM each route needs, transparent and multi-reference editing, the vendor benchmark, and the license terms you should read before shipping anything built on it. Every command below is copied from the official model card, the GitHub README, or the maintainer's README for that tool.

What is Qwen-Image-2.1?

Qwen-Image-2.1 is a unified image generation and editing model from the Qwen team at Alibaba. According to the official GitHub README, it is a single-stream diffusion transformer (DiT) built from three parts:

  • Transformer (the generator): 7B parameters across 32 layers, with block-causal attention. Text tokens use a causal mask, and image tokens use a chunk-level bidirectional mask.
  • Text encoder: Qwen3-VL 8B, a vision-language model. It encodes both your instruction and any reference images into one representation, which is how editing and generation share a single checkpoint.
  • VAE: a 64-channel RGBA autoencoder with 16x spatial compression. This is what gives the model a real alpha channel.
  • Scheduler: Flow Matching with Euler discrete sampling and dynamic shifting.

A note on the layer count: 36Kr's launch coverage describes a "20-layer Single-Stream DiT". The Hugging Face card and the GitHub README both say 32 layers, and the stable-diffusion.cpp docs refer to "the default 32-layer model". We go with 32.

The design choice that matters most for local users is prefix KV cache reuse. Your text and reference images are encoded once at the first denoising step and then reused for every later step. That is why multi-reference editing is much cheaper than you would expect for a model that accepts 10 input images.

Qwen-Image-2.1 key specs

SpecValueSource
Release dateSeptember 20, 2026 (HF repo created Sep 14)GitHub README news, HF API
Generator7B params, 32 single-stream DiT layersHF model card
Text encoderQwen3-VL 8BGitHub README
VAE64-channel RGBA, 16x spatial compressionGitHub README
Native resolution2048x2048 (presets up to 2752x1536 / 1536x2752)HF model card
Reference imagesUp to 10HF model card
TransparencyNative RGBA output and transparent-layer editingHF model card
Local editsCircles, painted annotations, or separate masksHF model card
Default steps40GitHub README
LicenseQwen Research License (non-commercial)LICENSE file

Qwen also ships two optional prompt enhancers: Qwen-Image-2.1-PE-T2I for text-to-image and Qwen-Image-2.1-PE-I2I for editing. Both are fine-tuned Qwen3.5-VL 9B checkpoints that turn a short prompt into a long, detailed description and suggest an aspect ratio. If you already run Qwen language models locally, our Qwen 3.8 lineup overview covers how the rest of the family fits together.

How much VRAM does Qwen-Image-2.1 need?

Qwen does not publish a VRAM table, so the numbers below come from the actual file sizes in each Hugging Face repo, plus the hardware guidance in Unsloth's run guide and one community test on an 8GB laptop GPU. Remember that you always need three pieces: the denoiser, the Qwen3-VL text encoder, and the VAE. The text encoder is often bigger than the denoiser.

SetupDenoiser fileText encoder fileVAERealistic GPU
Diffusers BF16 (official)~14.2 GB~17.5 GB~1.35 GB40GB+ fully on GPU; 24GB with enable_model_cpu_offload()
ComfyUI BF1614.23 GB17.53 GB (bf16)0.68 GB24GB+ with ComfyUI's automatic offloading
ComfyUI INT8 convrot (template default)7.26 GB9.35 GB (int8) or 6.31 GB (w4a8)0.68 GB16-24GB
GGUF Q8_0 (Unsloth)7.64 GB~5.2 GB (Q4 GGUF) or 6.31 GB (w4a8)0.68 GB12-16GB
GGUF Q6_K (Unsloth)6.27 GBsame0.68 GB12GB
GGUF Q5_K_M (Unsloth)5.39 GBsame0.68 GB8-12GB
GGUF Q4_K_M (Unsloth)4.20 GBsame0.68 GB8GB (community-tested on RTX 4060 Laptop)
GGUF Q3_K_M / Q2_K3.17 / 2.47 GBsame0.68 GB6-8GB, with visible quality loss

File sizes are decimal GB as listed by the Hugging Face API. ComfyUI's workflow notes list the same files in GiB, so its 6.76 GB for the INT8 denoiser is the same 7.26 GB file. The "Realistic GPU" column is our estimate, not a vendor figure. It assumes 1024x1024, batch size 1, and the text encoder being swapped out of VRAM once it has encoded your prompt. Native 2048x2048 output and multi-image edits need more headroom.

Unsloth's guide recommends GGUF Q4_K_M at 1024x1024 for 12-16GB cards. It also says FP8 can run on just 6GB of VRAM with offloading, at less than 2x slower, and that CPU-only machines with 12-16GB of RAM can run Q4_K_M with the Q4_K_XL text encoder. Unsloth defaults to INT8 over FP8 because INT8 scored the lower (better) LPIPS in its tests. Separately, the gatuwo multi-queue workflow repo reports running Q4_K_M with the INT8 text encoder on an 8GB RTX 4060 Laptop GPU.

How to run Qwen-Image-2.1 in ComfyUI (Method 1)

ComfyUI has supported Qwen-Image-2.1 natively since day one, with official workflow templates. This is the route we recommend for most people with a 16GB or 24GB card. If you have not set up ComfyUI before, our LTX-2 on ComfyUI guide walks through the base install.

Step 1: Update ComfyUI

Comfy's launch page says to update ComfyUI to the latest version before loading the weights, and the workflow notes point out that Desktop and Cloud builds follow stable releases, so a model supported in nightly may not show up there yet. The GGUF workflow repo we cite below lists ComfyUI 0.37.0 or newer. If the template or the TextEncodeQwenImage21 node is missing, your ComfyUI is too old.

Step 2: Download the files

Grab these from Comfy-Org/Qwen-Image-2.1. The official text-to-image template uses the INT8 files by default:

ComfyUI/
└── models/
    ├── diffusion_models/
    │   └── qwen_image_2.1_int8_convrot.safetensors      (or qwen_image_2.1_bf16.safetensors)
    ├── text_encoders/
    │   ├── qwen3vl_8b_int8_convrot.safetensors          (or qwen3vl_8b_bf16 / qwen3vl_8b_w4a8)
    │   └── qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors   (optional prompt enhancer)
    └── vae/
        └── qwen_image_2.1_vae_bf16.safetensors

The repo also includes qwen3.5_9b_qwen_image_2.1_pe_i2i.int8_convrot.safetensors for the editing prompt enhancer, plus Fun ControlNet Union patches under model_patches/. If VRAM is tight, pick qwen3vl_8b_w4a8.safetensors (6.31 GB) for the text encoder.

Step 3: Load the template

Open the template browser and load Qwen Image 2.1: Text to Image. There are also Image Edit and Remove Background templates. The same JSON files are on GitHub in Comfy-Org/workflow_templates if you prefer to drag them in.

Step 4: Use the right settings

The template's KSampler starts at 25 steps, CFG 1, euler sampler, simple scheduler, at 1024x1024. The workflow notes spell out the reasoning:

  • CFG: keep it at 1 for the official path. The negative prompt is ignored at CFG 1, so only raise CFG if you actually want a negative prompt.
  • Steps: the official pipeline uses about 40-50 with euler. The template starts at 25 for speed.
  • Resolution: default is 1 megapixel. For native 2K, choose 1:1 at 4 megapixels (2048x2048). Keep dimensions a multiple of 32.
  • Prompt enhancer: off by default. If you turn refine_prompt on, turn thinking_mode on too, because the enhancer checkpoint loses quality without its reasoning block.

The template also includes an experimental QwenImage21Cache node. Comfy's page says it lets you put the KV cache on auto, GPU, CPU or off. Storing it as int8 halves the cache "at about bf16 accuracy", and int4 quarters it but, per Comfy, "roughly doubles per-step error". That setting is worth changing on 12-16GB cards doing multi-reference edits.

How to run Qwen-Image-2.1 GGUF in ComfyUI on 8-12GB GPUs (Method 2)

GGUF quantizes only the denoiser. You still load the VAE and a Qwen3-VL text encoder separately. The two main sources are unsloth/Qwen-Image-2.1-GGUF, which uses Unsloth Dynamic 2.0 and upcasts sensitive layers per tensor, and leejet/Qwen-Image-2.1-GGUF, converted by the stable-diffusion.cpp author.

Step 1: Install the leejet fork of ComfyUI-GGUF

The original city96 ComfyUI-GGUF node does not recognize this architecture, and users report an "Unknown model architecture!" error. leejet's model card says to "use the leejet version of ComfyUI-GGUF, rather than the version maintained by city96". Install it the same way the ComfyUI-GGUF README describes, pointing at the leejet repo:

cd ComfyUI/custom_nodes
git clone https://github.com/leejet/ComfyUI-GGUF
pip install --upgrade gguf

If you already have city96's version in custom_nodes/ComfyUI-GGUF, remove or rename it first so the two don't clash. On the Windows portable build, install the requirements with the embedded Python as the README shows: .\python_embeded\python.exe -s -m pip install -r .\ComfyUI\custom_nodes\ComfyUI-GGUF\requirements.txt.

Step 2: Pick a quant

  • 8GB VRAM: Q4_K_M (4.20 GB). This has been tested on an RTX 4060 Laptop.
  • 10-12GB: Q5_K_M (5.39 GB) or Q6_K (6.27 GB).
  • 16GB: Q8_0 (7.64 GB), about the same size as the official INT8 file (7.26 GB), which you could use instead.

These are the file names and sizes in Unsloth's repo (for example qwen-image-2.1-Q4_K_M.gguf). leejet's repo names its files differently and has fewer rungs: qwen_image_2.1-Q4_K.gguf (4.20 GB), Q5_0 (5.07 GB), Q6_K (6.00 GB) and Q8_0 (7.69 GB), among others.

Put the .gguf file in ComfyUI/models/unet, the folder the ComfyUI-GGUF README specifies. Keep the VAE and text encoder in models/vae and models/text_encoders as in Method 1.

Step 3: Swap the loader

Load the official text-to-image template, then replace the stock "Load Diffusion Model" node with Unet Loader (GGUF), which the README lists under the bootleg category. Everything else stays the same. leejet also publishes a ready-made qwen_image_2_1_t2i_gguf.json workflow in the GGUF repo. For the text encoder, use qwen3vl_8b_w4a8.safetensors or qwen3vl_8b_int8_convrot.safetensors. The ComfyUI-GGUF node pack also includes GGUF-capable CLIP loader nodes if you would rather use a GGUF text encoder.

Alternative: stable-diffusion.cpp on the command line

If you want to skip ComfyUI, Unsloth's model card gives this sd-cli command. It uses Unsloth's Q4_K_XL Qwen3-VL text encoder and the VAE from unsloth/Qwen-Image-2.1-FP8:

sd-cli --diffusion-model qwen-image-2.1-Q4_K_M.gguf \
  --vae qwen_image_2.1_vae_bf16.safetensors \
  --llm Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf \
  -p "a cartoon sloth mascot waving, flat vector illustration, bright colours" \
  --steps 20 --cfg-scale 6.0 --sampling-method euler -W 1024 -H 1024 --diffusion-fa \
  -o out.png

Notice that sd.cpp uses CFG 6.0 while the Diffusers and ComfyUI paths use 1.0. Don't carry settings from one runtime to another. The stable-diffusion.cpp docs also add --offload-to-cpu for low-memory machines, and for editing with a GGUF text encoder you need the mmproj vision file, passed with --llm_vision. The docs also warn that the older Qwen-Image and Wan 2.2 VAEs do not work with 2.1, so use qwen_image_2.1_vae_bf16.safetensors.

How to run Qwen-Image-2.1 with Diffusers in Python (Method 3)

Diffusers added a dedicated QwenImage21Pipeline on launch day, and the GitHub README calls it the recommended integration. At launch it needed Diffusers from source. Install exactly as the model card shows:

pip install torch>=2.4.0
pip install transformers>=5.17
pip install git+https://github.com/huggingface/diffusers
pip install accelerate pillow

Text-to-image at native 2K, from the official card:

import torch
from diffusers import QwenImage21Pipeline

pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")

image = pipe(
    prompt="A neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night, reflections on wet pavement",
    width=2048, height=2048,
    num_inference_steps=40,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]

image.save("t2i_example.png")

At full BF16, the transformer plus text encoder add up to roughly 33GB of weights, so on a 24GB card use the card's memory optimization. Drop the .to("cuda") and enable CPU offload:

pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()

The recommended sizes from the README are 1:1 (2048x2048), 4:3 (2400x1792), 3:4 (1792x2400), 3:2 (2528x1696), 2:3 (1696x2528), 16:9 (2752x1536) and 9:16 (1536x2752).

Serving it as an API with vLLM-Omni

For a shared team endpoint rather than a desktop app, the README documents an OpenAI-style images endpoint through vLLM-Omni:

vllm serve Qwen/Qwen-Image-2.1 --omni --port 8091
curl http://localhost:8091/v1/images/generations \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen-Image-2.1",
    "prompt": "A ceramic teapot on a wooden table",
    "size": "1024x1024",
    "num_inference_steps": 40,
    "seed": 42
  }'

vLLM-Omni adds FP8 quantization, prefix KV caching, CUDA Graph decode and tensor parallelism. SGLang-Diffusion (sglang generate --model-path Qwen/Qwen-Image-2.1 ...) and LightX2V are the other supported servers. For the broader picture on running models behind your own endpoints, see our self-hosting LLMs guide.

How do you edit images and make transparent PNGs with Qwen-Image-2.1?

Single-image editing

Pass an input image and describe the change in plain language:

from PIL import Image

input_image = Image.open("input.png")

image = pipe(
    prompt="Change the background to a sunset beach",
    image=input_image,
    num_inference_steps=40,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]

image.save("edit_example.png")

Multi-reference composition (up to 10 images)

images = [Image.open(f"ref_{i}.png") for i in range(3)]
result = pipe(
    prompt="These three characters are sitting around a campfire in a forest",
    image=images,
    num_inference_steps=40,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
result.save("multi_ref_example.png")

Qwen's showcase includes a group photo built from six portrait references and a complete outfit assembled from five reference images (a model plus clothing, shoes, bag and hat). There is a small conflict here: Comfy's launch page says its text-encode node accepts up to 16 reference images, while Qwen's own card says up to 10. Treat 10 as the tested limit.

Transparent (RGBA) output

There is no transparency switch. The prompt decides whether the model outputs an alpha channel. Qwen's recommended format is:

This is an RGBA image with transparency. <your description>. The image has alpha channel and the background is transparent.

Save as PNG (or WebP). The stable-diffusion.cpp docs point out that saving to JPG drops the alpha channel. The same prompt format works for editing, which is how you cut a subject out of a photograph. ComfyUI's "Remove Background" template is built on the same idea. For local, targeted edits, the model also accepts circles, painted annotations or separate masks drawn on the input to mark the region to change.

Qwen-Image-2.1 benchmarks: how does it compare?

Independent benchmarks are thin so far. The main number in circulation is Qwen's own Qwen-Image-Bench score, reported in 36Kr's launch coverage:

ModelQwen-Image-Bench scoreOpen weights?Reported by
Qwen-Image-2.160.28Yes (research license)Qwen (vendor), via 36Kr
Nano Banana 2.059.82NoQwen (vendor), via 36Kr
GPT Image 1.559.65NoQwen (vendor), via 36Kr

Two caveats. First, this is a vendor-run benchmark that Qwen designed, so a 0.46-point lead over Nano Banana 2.0 is best read as "roughly on par with the top closed models", not as a clear win. Second, we could not find these scores on the Hugging Face card or in the GitHub README. Wait for LMArena or Artificial Analysis image leaderboards before making any firm quality claim. What holds up in hands-on use is the feature set: native RGBA, 10-image references and 2K output are hard to get from any other open-weight model right now. For a lighter, faster local option, compare it with our Z-Image Turbo install guide.

Can you use Qwen-Image-2.1 commercially?

No, not under the default license. Unlike the Apache-licensed Qwen language models, Qwen-Image-2.1 ships under the Qwen Research License Agreement, dated September 20, 2026. The key clauses:

  • You may use, modify and distribute the model "FOR NON-COMMERCIAL PURPOSES ONLY", and non-commercial is defined as "for research or evaluation purposes only".
  • Commercial use needs a separate license. Qwen directs requests to model-business@notice.qwencloud.com.
  • If you distribute copies, you must include a Notice file with Qwen's attribution text.
  • If you use the model or its outputs to train, fine-tune or improve an AI model that you distribute, you must prominently display "Built with Qwen" or "Improved using Qwen".

The GGUF and ComfyUI repackages carry the same license, so quantizing the model doesn't change anything. In practice, local experiments, research and internal evaluation are fine. Generating assets for a client, a product or paid marketing is not, unless you have the commercial license. This is not legal advice, so read the full LICENSE file before you ship anything.

Troubleshooting common Qwen-Image-2.1 errors

  • "Unknown model architecture!" when loading a GGUF: you are on city96's ComfyUI-GGUF. Replace it with leejet's fork.
  • ImportError: cannot import name 'QwenImage21Pipeline': your Diffusers release predates the pipeline. Install from source with pip install git+https://github.com/huggingface/diffusers, and make sure transformers>=5.17.
  • Template or TextEncodeQwenImage21 node missing: update ComfyUI. Desktop builds lag behind nightly.
  • Out of memory at 2048x2048: drop to 1024x1024 first, switch the text encoder to w4a8, set the Qwen Image 2.1 Cache node to CPU or int8, or step down one GGUF quant.
  • Washed-out or garbled images with a GGUF: check that you are using qwen_image_2.1_vae_bf16.safetensors. The original Qwen-Image VAE is not interchangeable.
  • No transparency in the output: use the exact RGBA prompt wording and save as PNG, not JPG.
  • Prompt enhancer output is poor: turn thinking_mode on together with refine_prompt.

Who should use Qwen-Image-2.1?

  • 24GB card (RTX 3090/4090/5090): use ComfyUI with the INT8 files, or BF16 if you want the reference output. This is the best overall experience.
  • 16GB card: use the ComfyUI INT8 denoiser with the w4a8 text encoder, or GGUF Q8_0.
  • 8-12GB card: use GGUF Q4_K_M to Q6_K through leejet's ComfyUI-GGUF, and stay at 1024x1024.
  • Developers building pipelines: use Diffusers for scripts, and vLLM-Omni or SGLang for a served endpoint.
  • Anyone doing commercial work: evaluate locally, but budget for the commercial license or pick a model with a permissive license.

If you are also running Qwen language models locally, the same hardware planning applies. See how to run Qwen 3.8 locally.

FAQ

What is Qwen Image 2.1?

Qwen-Image-2.1 is an open-weight image model from Alibaba's Qwen team, released September 20, 2026. It combines text-to-image generation and image editing in one checkpoint, with a 7B generator, a Qwen3-VL 8B text encoder, native transparent output and support for up to 10 reference images.

How much VRAM does Qwen Image 2.1 need?

Full BF16 in Diffusers needs about 33GB of weights, so a 24GB GPU needs CPU offload. The ComfyUI INT8 files suit 16-24GB. Community GGUF quants bring it down to 8GB, with Q4_K_M tested on an RTX 4060 Laptop GPU.

Is there a Qwen Image 2.1 GGUF?

Yes. Unsloth publishes GGUF quants from Q2_K (2.47 GB) up to F16 (14.23 GB), and leejet publishes Q2_K (2.56 GB) up to Q8_0 (7.69 GB). Use them in ComfyUI with the leejet fork of ComfyUI-GGUF, or with stable-diffusion.cpp. You still need the VAE and a Qwen3-VL text encoder.

Does Qwen-Image-2.1 work in ComfyUI?

Yes, natively from day one. Download the files from Comfy-Org/Qwen-Image-2.1 and load the official Text to Image, Image Edit or Remove Background templates after updating ComfyUI.

Can Qwen Image 2.1 generate transparent images?

Yes. Start the prompt with "This is an RGBA image with transparency." and end it with "The image has alpha channel and the background is transparent." Save as PNG to keep the alpha channel.

Is Qwen-Image-2.1 free for commercial use?

No. It uses the Qwen Research License, which allows research and evaluation only. Commercial use requires a separate license from Qwen, and if you use it or its outputs to train a model you distribute, you must display "Built with Qwen" or "Improved using Qwen".

Is Qwen-Image-2.1 better than GPT Image 1.5 or Nano Banana 2.0?

On Qwen's own Qwen-Image-Bench it scores 60.28, against 59.82 for Nano Banana 2.0 and 59.65 for GPT Image 1.5. That is a vendor-run benchmark with a very small margin, so treat it as roughly on par until independent leaderboards weigh in.

Sources

If your team is building products on open image and language models, Codersera can help you hire vetted remote developers who have shipped this kind of stack.