Overview
Qwen-Image 3.0 is Alibaba's third-generation image generation and editing foundation model, announced on July 21, 2026. The Qwen team's launch post compresses the pitch into one character โ ๅฎ, "Real" โ and splits it into three claims: "Rich Content" (prompts up to 4.5k tokens, for newspapers, exam papers and multi-panel layouts), "Authentic Details" (text legible at 10px, plus skin and hair micro-texture), and "Deep Knowledge" (native rendering of 12 languages and simulation of interfaces such as web pages, games and livestreams). Alibaba Cloud's community repost of the announcement is dated July 22, 2026.
It launched as an invite-only API preview. General availability came in early August 2026: Chinese tech press dated the full rollout to August 5, 2026, and OpenRouter lists qwen-image-3.0-pro as added on August 5 with Alibaba Cloud International as the sole provider.
The most consequential fact about the release is again not a benchmark. Qwen-Image 3.0 shipped with no weights, no licence, no model card, no technical report and no benchmark table. Hugging Face's API lists seven Qwen image repositories on September 6, 2026 and none of them is a 3.0; ModelScope's API answers record not found for Qwen/Qwen-Image-3.0 and Qwen/Qwen-Image-3.0-Pro while answering success for Qwen/Qwen-Image. That makes two closed flagships in a row, and it settles what was still an open question in February: Qwen-Image 2.0's closure was the new policy, not an exception.
This is the same trajectory the language models took. Qwen3.7-Max broke a long Apache 2.0 tradition by shipping as a proprietary API product while the sub-flagship line stayed open. One modality over, the openly licensed image checkpoints โ Qwen-Image-2512, Qwen-Image-Edit-2511 โ are now two generations behind and have not been refreshed since May 2026.
The prior generations, in short
Both are still relevant, because both are still what you can actually download or reason about.
- Qwen-Image (August 2025) โ a 20B MMDiT foundation model, Apache 2.0, with a same-day technical report (arXiv 2508.02324) carrying GenEval, DPG-Bench and OneIG-Bench tables. It is the last Qwen image model to publish an automated benchmark suite.
- Qwen-Image 2.0 (February 10, 2026) โ the first closed flagship. It unified generation and editing into one "omni-capable" model, coupling a Qwen3-VL encoder with an MMDiT backbone, added native 2K output and instructions "of up to 1K tokens", and got a technical report three months late (arXiv 2605.10730, May 11, 2026). Alibaba never published its parameter count; the "7B" figure that circulates for it comes from API resellers, not from Alibaba, and appears nowhere in the paper.
Capabilities
Drawn from Alibaba's launch post and Model Studio's 3.0 API reference. All of it is vendor-stated; none of it is independently measured.
- 4,500-token prompts โ the documented recommended maximum, against 1,300 tokens enforced by the 2.0 API and "up to 1K tokens" claimed in the 2.0 paper. This is the one specification that changed by a large factor, and it is what the "newspaper page, exam paper, storyboard in a single pass" demos rest on.
- 10px legible text โ the vendor claim is that labels, footnotes, LaTeX-style formulas and handwritten annotations stay readable at roughly ten pixels of glyph height, which is the size at which infographic captions actually live.
- 12 languages and 20+ fonts natively โ the Qwen line's long-standing strength was Chinese glyph fidelity; 3.0 extends the claim to a dozen scripts and to typeface control.
- Interface and document mimicry โ Alibaba demonstrates web pages, game HUDs, livestream overlays and knowledge diagrams, and claims the model can pull current world knowledge (for example a weather graphic for a named city and date) rather than inventing plausible-looking chrome.
- Unified generation and editing โ one model for both, with 1โ3 reference images accepted for instruction-driven edits, so there is no pipeline switch between a generator and an editor.
- Batch output โ the API's
nparameter accepts 1โ6 images per request, default 1, and unspecified sizes are auto-selected by the model from the prompt.
Technical Specifications
What Alibaba publishes about 3.0 is an API surface, not a model. Fields Alibaba does not publish are marked as such rather than estimated.
- Model IDs:
qwen-image-3.0-pro,qwen-image-3.0 - Dated snapshots: none published. The 2.0 series exposed pinnable snapshots (
qwen-image-2.0-pro-2026-06-22and earlier); the 3.0 aliases are unversioned - Prompt limit: "Recommended maximum: 4,500 tokens" for the positive prompt (tokens, not characters)
- Resolution: total output pixels between 512x512 and 2048x2048, aspect ratios 1:8 to 8:1; if
sizeis omitted the model picks a resolution from the prompt - Billing tiers: 1K and 2K, priced separately for the Pro model
- Output format: PNG. Images per call: 1โ6. Editing inputs: 1โ3 reference images
- Rate limits: 5 RPM for
qwen-image-3.0-pro, 20 RPM forqwen-image-3.0, in both the Beijing and Singapore regions - Parameters: not published. Architecture: not published for 3.0 โ the 2.0 report describes a Qwen3-VL encoder over a Multimodal Diffusion Transformer, and Alibaba has not said whether 3.0 keeps it. Training data: not published
- Weights: not published โ closed, API-only
- Platforms: Alibaba Cloud Model Studio (Bailian) / DashScope, Qwen Cloud, Qwen Chat, and OpenRouter via Alibaba Cloud International
The open-weight line
Seven Qwen image repositories remain on Hugging Face under Apache 2.0; the seventh, Qwen/Qwen-Image-Bench, is an evaluation model published May 21, 2026. Download counts are 30-day figures pulled from the Hugging Face API on September 6, 2026. Nothing has been added to this list since.
| Repository | Released | Task | Downloads (30d) |
|---|---|---|---|
Qwen/Qwen-Image-Edit-2509 | 2025-09-22 | Editing | 457,907 |
Qwen/Qwen-Image | 2025-08-04 | Text-to-image | 308,023 |
Qwen/Qwen-Image-Edit-2511 | 2025-12-23 | Editing | 286,809 |
Qwen/Qwen-Image-Edit | 2025-08-18 | Editing | 130,335 |
Qwen/Qwen-Image-2512 | 2025-12-31 | Text-to-image | 77,413 |
Qwen/Qwen-Image-Layered | 2025-12-19 | Layer decomposition | 76,475 |
Two Alibaba image lines, two teams
Alibaba ships image models from two separate groups, and they are easy to confuse. The Qwen team builds Qwen-Image. The Tongyi-MAI group builds Z-Image โ and Tongyi-MAI/Z-Image-Turbo, a genuinely Apache 2.0 6B model, recorded 677,373 downloads in the 30 days to September 6, 2026, more than any Qwen-Image checkpoint. Distribution and leaderboard placement still point in opposite directions: Qwen-Image 3.0 outranks Z-Image on arena.ai, and Z-Image outdownloads it by infinity, because you cannot download Qwen-Image 3.0 at all.
Use Cases
- Dense information graphics โ the 4,500-token prompt path is the reason to pick this model: a full infographic, exam paper or newspaper page laid out in one pass rather than assembled from crops.
- Slides, posters and storyboards โ multi-panel output where each cell has to carry its own legible caption.
- Multilingual marketing creative โ text-in-image across 12 scripts with typeface control, the Qwen line's durable advantage over Western image models.
- UI and product concepting โ web pages, dashboards and game overlays rendered as convincing mockups, when a designer needs a look rather than a shippable asset.
- Instruction-driven photo editing โ generation and editing in one model, 1โ3 reference images, no second pipeline.
- Self-hosted or air-gapped work โ not with 3.0. Use
Qwen-Image-2512,Qwen-Image-Edit-2511or Z-Image instead.
Performance / Benchmarks
Alibaba published no benchmark table for Qwen-Image 3.0 โ no GenEval, no DPG-Bench, no OneIG-Bench, no human-preference study, and no result on Alibaba's own Qwen-Image-Bench suite (arXiv 2605.28091). The launch is carried entirely by curated sample images. Everything quantitative below comes from third parties.
arena.ai text-to-image (retrieved September 6, 2026)
| Rank | Model | Elo | Votes | Arena's licence label |
|---|---|---|---|---|
| #1 | gpt-image-2 (medium) | 1382ยฑ4 | 77,830 | Proprietary |
| #8 | seedream-5.0-pro | 1258ยฑ4 | 59,484 | Proprietary |
| #9 | qwen-image-3.0-pro | 1254ยฑ7 | 10,923 | Proprietary |
| #16 | qwen-image-2.0-pro-2026-06-22 | 1191ยฑ6 | 12,016 | Proprietary |
| #39 | qwen-image-2512 | โ | โ | Apache 2.0 |
| #60 | qwen-image | โ | โ | Apache 2.0 |
Two things follow, both checkable at arena.ai/leaderboard/text-to-image:
- The generational gain is real: +63 Elo over 2.0-pro, and a jump from #16 to #9. That is a larger step than 2.0 made over the open checkpoints.
- "#1 among Chinese models" no longer holds. Chinese coverage at the August 5 rollout reported qwen-image-3.0-pro as fifth globally and first among Chinese models. A month later it is #9, and ByteDance's Seedream 5.0 Pro is one rank above it at #8. Arena rank claims decay in weeks; treat any quoted rank as a dated snapshot, including this one.
Artificial Analysis image arena (retrieved September 6, 2026)
A different organisation, a different Elo scale, and the two are not comparable with the arena.ai figures above. Artificial Analysis places Qwen-Image-3.0-Pro at #12 with an Elo of 1086 (5,551 samples) and Qwen-Image-3.0 at #14 with 1078 (5,177), against GPT Image 2 (high) at #1 with 1178.
Image editing
There is no Qwen-Image 3.0 entry on arena.ai's image-edit board as of September 6, 2026, even though the model is sold as an editor. The highest-placed Qwen editor there is still qwen-image-2.0-pro-2026-06-22 at #19 (1304ยฑ5, 50,824 votes), and the open qwen-image-edit sits at #31 with 1,981,112 votes โ roughly 39x the flagship's, and the most of any Apache 2.0 model on that board.
Limitations
- Closed weights, again. No self-hosting, no fine-tuning on your own hardware, no offline or air-gapped inference, no LoRA ecosystem. Two generations in, this is policy rather than an exception.
- No technical report and no model card. 2.0 at least produced a paper three months after launch. For 3.0 there is no architecture, no parameter count, no training description and no evaluation methodology โ you cannot reason about cost, latency or capability headroom from first principles.
- No vendor benchmarks at all. Every number on this page was produced by someone other than Alibaba, on preference arenas rather than task benchmarks.
- The 2K ceiling did not move. Total output pixels still cap at 2048x2048, the same as 2.0. Seedream 5.0's API reaches 4096x4096.
- No pinnable snapshots.
qwen-image-3.0-prois an unversioned alias. The 2.0 series let you pin a dated build; here the model behind the ID can change without a name change, which makes your own evaluations expire silently. - Pro is rate-limited to 5 RPM on Alibaba's published limits, so batch workloads land on the cheaper standard tier or on a quota increase request.
- 2K doubles the Pro price. The 1K/2K tier split means the resolution you want is a 2x cost decision on Pro, where the standard model charges one flat rate.
- The 4,500-token limit is "recommended", not specified. Alibaba does not document what happens above it; the 2.0 series silently truncated over-long prompts.
- Thin arena sample. 10,923 votes on text-to-image is a fraction of the 150,000+ behind the Gemini and GPT entries it is ranked against, so the placement will move.
- Content policy constraints. As a model served primarily from Chinese-mainland infrastructure, generation is shaped by content rules that differ from those governing US-hosted models.
Pricing & Access
Alibaba bills image models per image, never per token, and 3.0 adds a 1K/2K resolution split and a charge for reference images that the 2.0 launch card did not have. The Beijing rate card is native CNY; Alibaba's international regions restate it in CNY at an applied exchange rate, which is why the Singapore column below is full of long decimals.
| Model | Beijing (ๅ /ๅผ ) | Singapore, as Alibaba quotes it (ๅ /ๅผ ) |
|---|---|---|
qwen-image-3.0-pro, 1K output | ยฅ0.25 | ยฅ0.299768 |
qwen-image-3.0-pro, 2K output | ยฅ0.5 | ยฅ0.562065 |
qwen-image-3.0, 1K or 2K output | ยฅ0.18 | ยฅ0.224826 |
| Reference image input, either model | ยฅ0.02 | ยฅ0.022483 |
qwen-image-2.0-pro output | ยฅ0.5 | โ |
qwen-image-2.0 output | ยฅ0.18 | โ |
In USD, the native quotes come from Qwen Cloud, the international surface Alibaba routes non-mainland developers to: $0.04 per 1K image and $0.075 per 2K image for Pro, $0.03 flat for the standard model, and $0.003 per input image. Those are the rates OpenRouter publishes for the Alibaba Cloud International provider, and they match Artificial Analysis's listing of $40 and $30 per 1,000 images. The round numbers are the native quotes; the long decimals in the table above are conversions.
A free quota is listed on the Chinese pricing page โ 10 images, input and output combined, valid for 90 days from activation. That is a much smaller allowance than the 100-image quota advertised during the 2.0 era, and Alibaba's regional pages have contradicted each other about free quotas before. Confirm it in your own console before relying on it.
Access routes: Alibaba Cloud Model Studio / Bailian and DashScope, Qwen Cloud, OpenRouter (Alibaba Cloud International), and Qwen Chat to try it interactively. There is no self-hosted path. Rates change โ check the pricing page before budgeting.
Ecosystem & Tools
- Qwen-Image Generation and Editing 3.0 API reference โ the only primary specification of 3.0: prompt ceiling, size range,
n, reference images - Model Studio model pricing and ็พ็ผๆจกๅ่ฎก่ดน โ the per-image rate cards, English and Chinese
- Qwen-Image-3.0 announcement (Alibaba Cloud blog) โ the "Real" framing and every capability claim on this page
- QwenLM/Qwen-Image on GitHub โ inference code for the open checkpoints only; 3.0 added nothing to it
- Qwen on Hugging Face โ the seven Apache 2.0 image checkpoints; no 2.0, no 3.0
- ModelScope โ Alibaba's own hub; also carries no 3.0 weights
- Qwen Chat โ the consumer surface where 3.0 is served
- Qwen-Image-Lightning and vLLM-Omni โ acceleration and serving stacks, built for the open checkpoints only
Community & Resources
- Qwen-Image-2.0 Technical Report (arXiv 2605.10730) - May 11, 2026; the last architecture disclosure the line produced
- Qwen-Image Technical Report (arXiv 2508.02324) - the original 20B MMDiT model and its benchmark tables
- Qwen-Image-Bench (arXiv 2605.28091) - Alibaba's own text-to-image evaluation suite, which carries no published 3.0 result
- Qwen-Image-Layered (arXiv 2512.15603) - layer decomposition for inherent editability
- arena.ai text-to-image leaderboard and image-edit leaderboard - the live Elo figures quoted above
- Artificial Analysis text-to-image leaderboard - a separate Elo scale, not comparable with arena.ai
- Qwen - the official site; note that its blog renders client-side and exposes no static text
- Compare with Seedream 5.0, Nano Banana, Z-Image, HunyuanImage 3.0, HiDream-O1-Image, and Stable Diffusion 3.5