Qwen3.8 Max-Prime and Omni-Flash: What Actually Changed

Alibaba added three Qwen3.8 SKUs in September with no launch post. Max-Prime is a higher-throughput serving tier at double the price, not new weights. Omni-Flash accepts audio and video but returns text only, at Flash pricing. Full lineup, pricing and licence tables.

Quick answer. Qwen3.8 Max-Prime is not a new model — it is Qwen3.8 Max served as a separate higher-throughput SKU at $4/$12 per 1M tokens, double Max's $2/$6. Qwen3.8 Omni-Flash is a genuinely new capability: it accepts text, image, audio and video but returns text only, at Flash pricing of $0.15/$0.47.

Between 3 and 23 September 2026, Alibaba quietly added three more SKUs to the Qwen 3.8 family: qwen3.8-max-0902, qwen3.8-omni-flash and qwen3.8-max-prime. None of them got a launch blog post. Two of them are not in Alibaba's public model catalogue at all.

That matters because the naming invites exactly the wrong conclusion. "Max-Prime" sounds like a tier above Max — new weights, higher scores, worth the premium. It is not. And "Omni-Flash" sounds like it should talk back in a voice. It does not. Here is what each one actually is, verified against Alibaba's own pricing pages, the Hugging Face model cards, and the live API listings.

What is Qwen3.8 Max-Prime, and why does it cost double?

Max-Prime appeared on OpenRouter on 23 September 2026 at $4 per million input tokens and $12 per million output tokens. Qwen3.8 Max and its dated snapshot qwen3.8-max-0902 both sit at $2 and $6. The cached-input rate doubles in lockstep too: $0.50 per million for Max-Prime against $0.25 for Max.

The only description Alibaba has published anywhere is the one carried on the API listing, and it is unusually blunt:

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video input and returns text, with a 1M-token context window and reasoning enabled by default. Tool calling, structured outputs, and configurable reasoning effort are supported, matching Qwen3.8 Max.

"Matching Qwen3.8 Max" is the operative phrase. Every spec is identical — 1,000,000-token context, 131,072-token maximum output, the same text/image/video input surface, the same reasoning_effort control, the same tool-calling and structured-output support. The only thing that differs is throughput and the price.

We looked for evidence of new weights and found none anywhere:

  • Max-Prime is absent from Alibaba Cloud Model Studio's English model catalogue and from the Chinese catalogue at help.aliyun.com.
  • Qwen Cloud, Alibaba's own first-party API storefront, lists Qwen3.8-Max-0902 and Qwen3.8-Flash but has no Max-Prime page — /models/qwen3.8-max-prime returns a 404.
  • No Hugging Face repository exists under the Qwen organisation for Max-Prime. The newest Qwen3.8 repo is Qwen3.8-Flash-Next, published 24 August 2026.
  • Qwen's launch post for the family, "Qwen3.8-Max: A New Bar for Coding and Cowork" (3 August 2026), does not mention Max-Prime, and no benchmark table has been published for it separately from Max.

So the honest read is this: Max-Prime is a serving tier, not a model. You are paying a 2× premium for capacity and speed on the same weights. If Alibaba later publishes distinct benchmarks, that changes — but as of early October 2026 nothing of the sort exists, and an exact 2× on input, output and cache alike is what you would expect from a capacity SKU rather than a smarter model. If your workload is not throughput-constrained, Max-Prime buys you nothing: point at qwen3.8-max, or pin qwen3.8-max-0902, and keep the other half of your budget.

What is Qwen3.8 Omni-Flash, and which modalities does it accept?

Omni-Flash, listed 21 September 2026, is the far more interesting release — and it is priced at $0.15 input and $0.47 output, identical to the text-only Qwen3.8 Flash. Alibaba describes it as "Qwen's next-generation native omni-modal model," and the architecture note explains how the price is possible: it is built on the Qwen3.8-Flash-Next architecture, the same sparse backbone that powers ordinary Flash.

Here is the modality list, which is the single most misunderstood thing about this model:

ModalityInputOutput
TextYesYes
ImageYesNo
VideoYesNo
Audio / speechYesNo

Audio in, text out. Three independent sources agree, which is worth spelling out because the "Omni" label has meant audio generation in every previous Qwen Omni release:

  • The API listing states plainly that "Qwen3.8 Omni Flash accepts text, images, audio, and video as input and returns text."
  • Alibaba Cloud's Model Studio documentation gives the signature as "Input: Text, images, audio, video | Output: Text."
  • Qwen Cloud's own code sample for the model carries the comment # Qwen3.8-Omni-Flash supports text output only. Do not set audio.

What you get for $0.15 per million tokens is therefore not a voice model — it is an audio and video comprehension model at commodity text pricing, with a 1M-token context window. That is still a notable thing. Alibaba calls out two-channel and four-channel spatial audio understanding, and positions the model for multimedia summarisation, audio-video dialogue, film and video production workflows, and narration, alongside the usual coding, knowledge work and GUI-interaction duties inherited from Flash.

One small operational difference from Flash: Omni-Flash offers implicit prompt caching at $0.016 per million tokens but no explicit cache-creation SKU, whereas Flash exposes explicit cache creation at $0.20 and reads at $0.016.

Can Omni-Flash speak back, or only write?

Only write. If you need generated speech, Alibaba sells that as a separate model: qwen3.8-omni-flash-realtime. It is a genuinely different product with a genuinely different billing shape — full-duplex audio-video interaction over realtime protocols, aimed at smart devices, robotics and interactive agents, with audio input in over 60 languages and speech output in over 30. It supports multichannel audio input, function calling and MCP integration.

Its pricing is per-modality rather than flat, and it is roughly 4× Omni-Flash on the output side:

Omni-Flash-Realtime billing linePrice per 1M tokens
Input: text$0.70
Input: audio$0.93
Input: text / image / video$0.23
Output: text and audio (output text not charged)$1.87

The rule is simple: Omni-Flash listens and reads; Omni-Flash-Realtime converses. Picking the wrong one is the likeliest integration mistake in this release.

What does the full Qwen 3.8 lineup look like now?

There are nine SKUs in play if you count the hosted and open variants separately. Prices below are Alibaba's own first-party list rates from Qwen Cloud; context figures are the hosted service's.

VariantParametersContextIn → outWeights & licence$/1M in → outBuilt for
Qwen3.8-27B27B dense + vision encoder262K native, to 1Mtext, image, video → textOpen, Apache 2.0$0.50 → $3.00Compact self-hosted VLM
Qwen3.8-Flash-Next125B total / 6B active, + 51B n-gram embedding + 4B MTP262K native, to 1Mtext, image, video → textOpen, Qwen Community License 1.0Self-host onlyQwen4-architecture preview
Qwen3.8-FlashHosted Flash-Next1Mtext, image, video → textClosed service$0.15 → $0.47High-volume multimodal work
Qwen3.8-Omni-FlashFlash-Next architecture1Mtext, image, audio, video → textClosed, API-only$0.15 → $0.47Audio + video comprehension
Qwen3.8-Omni-Flash-RealtimeFlash-Next architecture1Mtext, image, audio, video → text + audioClosed, API-onlyper-modality (see table above)Full-duplex voice agents
Qwen3.8-2.4T-A95B2.4T total / 95B active262K native, to 1,010,000text → textOpen, Qwen3.8-Max LicenseSelf-host onlyOpen-weight flagship
Qwen3.8-MaxHosted 2.4T-A95B + vision1Mtext, image, video → textClosed service$2.00 → $6.00Frontier reasoning and agents
Qwen3.8-Max-0902Dated Max snapshot1Mtext, image, video → textClosed service$2.00 → $6.00Pinned, reproducible builds
Qwen3.8-Max-PrimeSame weights as Max1Mtext, image, video → textClosed, API-only$4.00 → $12.00Higher-throughput capacity

Two relationships in that table are easy to miss and genuinely useful. Qwen3.8-Max is the hosted build of Qwen3.8-2.4T-A95B, with vision input, a non-thinking mode, 1M context by default and built-in tools added on top of the open weights. And Qwen3.8-Flash is the hosted build of Qwen3.8-Flash-Next. In both cases the open release is the same model minus the production conveniences — which means you can prototype against the API and move to self-hosting without changing models, a trick that does not work with most frontier vendors. We cover that migration path in our guide to running Qwen 3.8 locally.

Flash-Next deserves a footnote of its own. Qwen describes it as "an experimental preview of the architecture that will underpin Qwen4," introducing Qwen Sparse Attention at micro-block level, Gated Residual streams, and a 20-million-entry n-gram embedding table. At 125B total parameters with 6B activated, it is the cheapest route into frontier-class quality if you have the VRAM.

How do these models actually benchmark?

Here is the uncomfortable part, and we are keeping vendor claims and independent measurement strictly separate.

Neither Max-Prime nor Omni-Flash has any published benchmark results — vendor or independent. No Qwen blog post covers either model. Neither appears in Alibaba's own documentation with a score table. Neither is listed on Artificial Analysis (their dedicated model pages both return 404, and neither appears among the 224 models in its full Intelligence Index v4.3.2 field) nor on the LMArena text leaderboard. Any number you see quoted for Max-Prime is almost certainly Qwen3.8 Max's number — defensible if they are the same weights, but an inference rather than a measurement.

Independent measurements

Third-party evaluators have covered the rest of the family. These figures are independently measured, not vendor-supplied. Artificial Analysis figures are Intelligence Index v4.3.2 (October 2026) with the effort tier named; AA rebased from v4.1.1 to v4.3.2 and there is no conversion factor, so do not compare these against older scores. AA ranks come in two denominators — /224 for the whole field and /118 for open-weights models only — and the ranks below are the full /224 field:

Model (effort tier)AA Index v4.3.2Cost per Index taskOutput speedTime to first tokenLMArena Elo
Qwen3.8 Max (0902)45.42 (#32 / 224)$5.4137.15 tok/s2.64 s1482 (rank #23)
Qwen3.8 Max (0803, the undated qwen3.8-max)40.15$2.6737.30 tok/s2.63 s—
Qwen3.8 2.4T-A95B39.89$2.1637.72 tok/s2.75 s—
Qwen3.8-Flash-Next39.82$0.3755.23 tok/s2.48 s—
Qwen3.8 27B (Xhigh)33.70$1.0145.17 tok/s3.79 s1438 (rank #99)
Qwen3.8 Max-PrimeNot listed
Qwen3.8 Omni-FlashNot listed

Three things in that table are worth dwelling on. First, Flash-Next scores 39.82 against the 2.4T open flagship's 39.89 — a 0.07-point difference — while running faster and costing $0.37 per Index task against $2.16. The architectural work is real, not marketing, and the cost ratio is where it shows: Flash-Next completes an Index task for a fourteenth of what Max (0902) spends ($0.37 against $5.41) for 5.6 index points less. (Use that per-task column, not AA's whole-suite totals, for ratios — AA runs different task counts per model, so suite totals are not comparable across rows.)

Second, pinning the undated qwen3.8-max string costs you 5.27 index points for free. The 0902 snapshot scores 45.42 against 40.15 for the original 3 August weights, at the identical $2/$6 rate — the single cheapest upgrade available in this family is editing a model string. It does cost more per completed task ($5.41 against $2.67), because 0902 thinks more.

Third, and most revealing for our purposes: Qwen3.8 Max measures only 37.15 tokens per second, ranking roughly #180 of 224 models on output speed. Max is genuinely slow.

That reframes Max-Prime considerably. A throughput SKU is not an arbitrary upsell — it is a targeted response to a measured weakness in the product. It also means the thing you are buying with the 2× premium is the thing Max is actually worst at, which is a more coherent proposition than the naming suggests. It still is not more intelligence.

One further independent data point: Artificial Analysis scores Qwen3.8 27B at four reasoning_effort settings on v4.3.2 — 33.70 at xhigh, 27.55 at medium, 26.20 at low, 20.15 with reasoning off. A 13.6-point swing from one parameter, worth tuning before reaching for a larger model. Qwen3.8 Max is also the #3 model by weekly token usage on OpenRouter, so whatever the leaderboards say, it is being used in volume.

Vendor claims

Alibaba's own table for Qwen3.8 Max, from the 3 August 2026 launch post. These are vendor-reported figures, self-run:

Benchmark (vendor-reported)Qwen3.8-MaxOpus 4.8Fable 5GPT-5.6 Sol
Terminal Bench 2.186.684.684.688.8
SWE-bench Pro67.769.280.064.6
GPQA Diamond92.692.092.694.1
Humanity's Last Exam43.645.753.347.2
IFBench82.862.263.572.7
OSWorld-Verified86.183.485.083.2
VideoMME (w/ subs)90.485.4—89.5

For Omni-Flash, the nearest available proxy is the Flash-Next model card, since Omni-Flash is built on that architecture. Again, vendor-reported:

Benchmark (vendor-reported)Flash-NextQwen3.8-27BQwen3.7-PlusDeepSeek-V4-Flash-0731
SWE-bench Pro62.561.755.856.0
SWE-bench Multilingual81.073.875.8—
LiveCodeBench v691.990.389.690.6
GPQA Diamond91.789.290.390.8
Humanity's Last Exam35.930.834.733.8
Toolathlon Verified (Pass@1)73.567.150.670.3
LVBench (long video)76.672.476.263.0

Treat all of the above as directional. Self-reported tables are chosen by the vendor, and the gap between the Max tier and the Flash tier on these particular benchmarks is narrower than the roughly 13× difference in their hosted prices would suggest.

What do the Qwen 3.8 models cost, really?

One pricing caveat worth internalising. Public model aggregators often display the cheapest provider rather than the vendor's own rate. For Qwen3.8-27B that distortion is real: third-party hosts list it near $0.425 input and $2.55 output, below Alibaba's own $0.50 and $3.00, because the weights are Apache-licensed and anyone can serve them. For Max-Prime and Omni-Flash it does not apply — Alibaba is the sole provider of both, and the aggregator figures match Qwen Cloud's published rates to the cent.

Across the hosted range, all four main SKUs share one operational envelope: 1M context, 991K maximum input (983K with thinking on), 131K maximum output, 262K maximum reasoning tokens, and a 2M tokens-per-minute limit.

Hosted SKUInputOutputCached inputOutput vs Flash
Qwen3.8-Flash$0.15$0.47$0.0161×
Qwen3.8-Omni-Flash$0.15$0.47$0.0161×
Qwen3.8-27B$0.50$3.00$0.106.4×
Qwen3.8-Max / Max-0902$2.00$6.00$0.2512.8×
Qwen3.8-Max-Prime$4.00$12.00$0.5025.5×

The headline here is that Omni-Flash adds an entire input modality for free. Audio understanding at the same per-token rate as text-only Flash is the genuine bargain in this release, and it is the opposite of the Max-Prime story.

Are the weights open, and under which licence?

This is where the Qwen family gets genuinely treacherous, because three different licences are in play and "Qwen is open-source" is only partly true. Neither of the two new September models has published weights at all — both are API-only.

VariantWeightsLicenceSeparate licence needed for MaaS / AI coding assistant?
Qwen3.8-27BPublishedApache 2.0No
Qwen3.8-Flash-NextPublishedQwen Community License 1.0Yes — at any revenue
Qwen3.8-2.4T-A95BPublishedQwen3.8-Max LicenseYes, above $50M revenue in any 12 months
Qwen3.8-Max / Max-0902 / Max-PrimeNot publishedClosed service termsn/a
Qwen3.8-Flash / Omni-Flash / Omni-Flash-RealtimeNot publishedClosed service termsn/a

That fourth column is what catches companies out. Both community licences require a separate agreement from Qwen if you run a "Model as a Service" or "AI Work Assistant" business — giving third parties inference or fine-tuning access, or shipping a product primarily designed for AI-assisted coding or office productivity. But the thresholds differ, and counter-intuitively the smaller model is the stricter one:

  • Qwen3.8-Flash-Next (Qwen Community License 1.0) attaches no revenue floor. If you are in either of those businesses, you need a separate licence before any commercial use, full stop.
  • Qwen3.8-2.4T-A95B (Qwen3.8-Max License) only triggers the requirement once the licensee and affiliates exceed US$50,000,000 of aggregate revenue over any consecutive twelve months.

Both licences also carry an attribution clause: any product above 100 million monthly active users or US$20 million monthly revenue must display the model name prominently in its interface. Internal use is explicitly carved out of the MaaS restriction in both, provided you do not expose the model, its outputs or its capabilities to third parties.

The upshot: if you are building a coding tool or an inference product, Qwen3.8-27B under Apache 2.0 is the only variant you can ship without a lawyer. Check the LICENSE file yourself for anything else — Hugging Face reports both non-Apache variants simply as "other," which tells you nothing.

Which Qwen 3.8 model should you use?

A decision rule, in priority order:

  1. Do you need generated speech? Then qwen3.8-omni-flash-realtime, and nothing else in the family will do.
  2. Do your inputs include audio or video you need understood? Then qwen3.8-omni-flash. At Flash pricing there is no reason to reach higher.
  3. Do you need to ship weights in a commercial coding or inference product? Then Qwen3.8-27B, because Apache 2.0 is the only licence here without a separate-agreement clause.
  4. Are you self-hosting for internal use and want maximum quality per GPU? Then Qwen3.8-Flash-Next — 6B active parameters, frontier-adjacent scores.
  5. Do you need frontier reasoning through an API? Then qwen3.8-max-0902 — not the undated qwen3.8-max. Same $2/$6 rate, but Artificial Analysis measures the 0902 snapshot at 45.42 on Intelligence Index v4.3.2 against 40.15 for the undated endpoint. You get the pinned-snapshot reproducibility and the higher score in the same model string. Budget for a higher cost per completed task ($5.41 against $2.67), since 0902 spends more reasoning tokens.
  6. Are you actually throughput-constrained on Max, with the rate limit as your bottleneck rather than cost? Only then does qwen3.8-max-prime make sense.
Use casePickWhy
Meeting and call summarisationOmni-FlashAudio input at text prices, 1M context
Video review and editing workflowsOmni-FlashNative video plus spatial audio understanding
Voice assistant, robotics, smart deviceOmni-Flash-RealtimeOnly SKU that emits audio
High-volume text and vision classificationFlashCheapest hosted tier
Commercial product shipping weights27BApache 2.0, no separate licence
Long-horizon agents, hard reasoningMax-090245.42 on AA v4.3.2, best in the family
Reproducible evaluation harnessMax-0902Dated snapshot, same price, 5.27 AA points higher
Rate-limited at scale on MaxMax-PrimeCapacity, not capability

If you are coming at this cold, start with our Qwen 3.8 model lineup guide for the family overview, then the deep dives on Qwen3.8 Max, Qwen3.8 Flash and Qwen3.8 27B.

So what is the verdict?

September's three additions break into one real capability and two packaging exercises. Omni-Flash is the release worth your attention: audio and video comprehension at $0.15 per million input tokens with a 1M-token window, plus spatial-audio support. Max-0902 is housekeeping — a dated snapshot at the same price as Max.

Max-Prime is the one to be sceptical about. Until Alibaba publishes weights or benchmarks that distinguish it from Max, treat it as a capacity lane with a 2× price tag, and buy it only when throughput — not quality — is what you are short of.

FAQ

What is Qwen3.8 Max-Prime?

Qwen3.8 Max-Prime is a higher-throughput serving variant of Qwen3.8 Max, listed on 23 September 2026 at $4 per million input tokens and $12 per million output tokens. Alibaba's own description says it is "served as a separate SKU at a higher price point" with capabilities "matching Qwen3.8 Max." It has the same 1M context window, the same 131,072-token output cap, the same text/image/video input surface and the same tool-calling and reasoning-effort controls.

What is Qwen3.8 Omni-Flash?

Qwen3.8 Omni-Flash is Alibaba's omni-modal model listed on 21 September 2026, built on the Qwen3.8-Flash-Next architecture. It accepts text, images, audio and video as input and returns text, with a 1M-token context window, at $0.15 input and $0.47 output per million tokens. It supports two-channel and four-channel spatial audio understanding and is aimed at multimedia summarisation, audio-video dialogue and video production workflows.

Does Qwen3.8 Omni-Flash support audio?

It supports audio input but not audio output. Alibaba's documentation gives the signature as "Input: Text, images, audio, video | Output: Text," and Qwen Cloud's code sample carries the explicit warning that the model "supports text output only. Do not set audio." If you need generated speech, use the separate qwen3.8-omni-flash-realtime model, which does full-duplex audio conversation with speech output in over 30 languages.

How much do Qwen3.8 Max-Prime and Omni-Flash cost?

Max-Prime is $4.00 per million input tokens and $12.00 per million output tokens, with cached input at $0.50 — exactly double Qwen3.8 Max's $2.00/$6.00/$0.25. Omni-Flash is $0.15 input and $0.47 output with cached input at $0.016, identical to text-only Qwen3.8 Flash. Both figures are Alibaba's own first-party rates, since Alibaba is the sole provider of both models.

Are Qwen3.8 Max-Prime and Omni-Flash open source?

No. Neither has published weights; both are API-only. Within the wider family, Qwen3.8-27B is Apache 2.0, Qwen3.8-Flash-Next is under the Qwen Community License 1.0, and Qwen3.8-2.4T-A95B is under a separate Qwen3.8-Max License. Qwen3.8 Max, Max-0902, Max-Prime, Flash, Omni-Flash and Omni-Flash-Realtime are all closed hosted services.

Which Qwen 3.8 model should I use?

Use Omni-Flash if your inputs include audio or video, Omni-Flash-Realtime if you need spoken output, Flash for cheap high-volume text and vision work, Qwen3.8-27B if you need Apache-2.0 weights you can ship commercially, Flash-Next for the best quality per GPU when self-hosting internally, and Max for frontier reasoning through an API. Choose Max-Prime only when throughput rather than cost or quality is your constraint.

Is Qwen3.8 Max-Prime worth double the price of Qwen3.8 Max?

Only if you are throughput-constrained. There is no published evidence that Max-Prime has different weights or better quality than Max: it is absent from Alibaba's English and Chinese model catalogues and from Qwen Cloud's model list, has no Hugging Face repository, was not mentioned in the Qwen3.8-Max launch post, and has no benchmark table of its own. The pricing doubles uniformly across input, output and cached input, which is the signature of a capacity tier rather than a more capable model.