Developer
Alibaba

Qwen-Image 3.0

Alibaba's third-generation image model, announced July 21, 2026: 4,500-token prompts, 10px text, 12 languages โ€” and a second flagship with no weights.

Updated

Released
Jul 21, 2026
Type
Image Generation Model
License
Proprietary
On this page

Overview

Qwen-Image 3.0 is Alibaba's third-generation image generation and editing foundation model, announced on July 21, 2026. The Qwen team's launch post compresses the pitch into one character โ€” ๅฎž, "Real" โ€” and splits it into three claims: "Rich Content" (prompts up to 4.5k tokens, for newspapers, exam papers and multi-panel layouts), "Authentic Details" (text legible at 10px, plus skin and hair micro-texture), and "Deep Knowledge" (native rendering of 12 languages and simulation of interfaces such as web pages, games and livestreams). Alibaba Cloud's community repost of the announcement is dated July 22, 2026.

It launched as an invite-only API preview. General availability came in early August 2026: Chinese tech press dated the full rollout to August 5, 2026, and OpenRouter lists qwen-image-3.0-pro as added on August 5 with Alibaba Cloud International as the sole provider.

The most consequential fact about the release is again not a benchmark. Qwen-Image 3.0 shipped with no weights, no licence, no model card, no technical report and no benchmark table. Hugging Face's API lists seven Qwen image repositories on September 6, 2026 and none of them is a 3.0; ModelScope's API answers record not found for Qwen/Qwen-Image-3.0 and Qwen/Qwen-Image-3.0-Pro while answering success for Qwen/Qwen-Image. That makes two closed flagships in a row, and it settles what was still an open question in February: Qwen-Image 2.0's closure was the new policy, not an exception.

This is the same trajectory the language models took. Qwen3.7-Max broke a long Apache 2.0 tradition by shipping as a proprietary API product while the sub-flagship line stayed open. One modality over, the openly licensed image checkpoints โ€” Qwen-Image-2512, Qwen-Image-Edit-2511 โ€” are now two generations behind and have not been refreshed since May 2026.

The prior generations, in short

Both are still relevant, because both are still what you can actually download or reason about.

  • Qwen-Image (August 2025) โ€” a 20B MMDiT foundation model, Apache 2.0, with a same-day technical report (arXiv 2508.02324) carrying GenEval, DPG-Bench and OneIG-Bench tables. It is the last Qwen image model to publish an automated benchmark suite.
  • Qwen-Image 2.0 (February 10, 2026) โ€” the first closed flagship. It unified generation and editing into one "omni-capable" model, coupling a Qwen3-VL encoder with an MMDiT backbone, added native 2K output and instructions "of up to 1K tokens", and got a technical report three months late (arXiv 2605.10730, May 11, 2026). Alibaba never published its parameter count; the "7B" figure that circulates for it comes from API resellers, not from Alibaba, and appears nowhere in the paper.

Capabilities

Drawn from Alibaba's launch post and Model Studio's 3.0 API reference. All of it is vendor-stated; none of it is independently measured.

  • 4,500-token prompts โ€” the documented recommended maximum, against 1,300 tokens enforced by the 2.0 API and "up to 1K tokens" claimed in the 2.0 paper. This is the one specification that changed by a large factor, and it is what the "newspaper page, exam paper, storyboard in a single pass" demos rest on.
  • 10px legible text โ€” the vendor claim is that labels, footnotes, LaTeX-style formulas and handwritten annotations stay readable at roughly ten pixels of glyph height, which is the size at which infographic captions actually live.
  • 12 languages and 20+ fonts natively โ€” the Qwen line's long-standing strength was Chinese glyph fidelity; 3.0 extends the claim to a dozen scripts and to typeface control.
  • Interface and document mimicry โ€” Alibaba demonstrates web pages, game HUDs, livestream overlays and knowledge diagrams, and claims the model can pull current world knowledge (for example a weather graphic for a named city and date) rather than inventing plausible-looking chrome.
  • Unified generation and editing โ€” one model for both, with 1โ€“3 reference images accepted for instruction-driven edits, so there is no pipeline switch between a generator and an editor.
  • Batch output โ€” the API's n parameter accepts 1โ€“6 images per request, default 1, and unspecified sizes are auto-selected by the model from the prompt.

Technical Specifications

What Alibaba publishes about 3.0 is an API surface, not a model. Fields Alibaba does not publish are marked as such rather than estimated.

  • Model IDs: qwen-image-3.0-pro, qwen-image-3.0
  • Dated snapshots: none published. The 2.0 series exposed pinnable snapshots (qwen-image-2.0-pro-2026-06-22 and earlier); the 3.0 aliases are unversioned
  • Prompt limit: "Recommended maximum: 4,500 tokens" for the positive prompt (tokens, not characters)
  • Resolution: total output pixels between 512x512 and 2048x2048, aspect ratios 1:8 to 8:1; if size is omitted the model picks a resolution from the prompt
  • Billing tiers: 1K and 2K, priced separately for the Pro model
  • Output format: PNG. Images per call: 1โ€“6. Editing inputs: 1โ€“3 reference images
  • Rate limits: 5 RPM for qwen-image-3.0-pro, 20 RPM for qwen-image-3.0, in both the Beijing and Singapore regions
  • Parameters: not published. Architecture: not published for 3.0 โ€” the 2.0 report describes a Qwen3-VL encoder over a Multimodal Diffusion Transformer, and Alibaba has not said whether 3.0 keeps it. Training data: not published
  • Weights: not published โ€” closed, API-only
  • Platforms: Alibaba Cloud Model Studio (Bailian) / DashScope, Qwen Cloud, Qwen Chat, and OpenRouter via Alibaba Cloud International

The open-weight line

Seven Qwen image repositories remain on Hugging Face under Apache 2.0; the seventh, Qwen/Qwen-Image-Bench, is an evaluation model published May 21, 2026. Download counts are 30-day figures pulled from the Hugging Face API on September 6, 2026. Nothing has been added to this list since.

RepositoryReleasedTaskDownloads (30d)
Qwen/Qwen-Image-Edit-25092025-09-22Editing457,907
Qwen/Qwen-Image2025-08-04Text-to-image308,023
Qwen/Qwen-Image-Edit-25112025-12-23Editing286,809
Qwen/Qwen-Image-Edit2025-08-18Editing130,335
Qwen/Qwen-Image-25122025-12-31Text-to-image77,413
Qwen/Qwen-Image-Layered2025-12-19Layer decomposition76,475

Two Alibaba image lines, two teams

Alibaba ships image models from two separate groups, and they are easy to confuse. The Qwen team builds Qwen-Image. The Tongyi-MAI group builds Z-Image โ€” and Tongyi-MAI/Z-Image-Turbo, a genuinely Apache 2.0 6B model, recorded 677,373 downloads in the 30 days to September 6, 2026, more than any Qwen-Image checkpoint. Distribution and leaderboard placement still point in opposite directions: Qwen-Image 3.0 outranks Z-Image on arena.ai, and Z-Image outdownloads it by infinity, because you cannot download Qwen-Image 3.0 at all.

Use Cases

  • Dense information graphics โ€” the 4,500-token prompt path is the reason to pick this model: a full infographic, exam paper or newspaper page laid out in one pass rather than assembled from crops.
  • Slides, posters and storyboards โ€” multi-panel output where each cell has to carry its own legible caption.
  • Multilingual marketing creative โ€” text-in-image across 12 scripts with typeface control, the Qwen line's durable advantage over Western image models.
  • UI and product concepting โ€” web pages, dashboards and game overlays rendered as convincing mockups, when a designer needs a look rather than a shippable asset.
  • Instruction-driven photo editing โ€” generation and editing in one model, 1โ€“3 reference images, no second pipeline.
  • Self-hosted or air-gapped work โ€” not with 3.0. Use Qwen-Image-2512, Qwen-Image-Edit-2511 or Z-Image instead.

Performance / Benchmarks

Alibaba published no benchmark table for Qwen-Image 3.0 โ€” no GenEval, no DPG-Bench, no OneIG-Bench, no human-preference study, and no result on Alibaba's own Qwen-Image-Bench suite (arXiv 2605.28091). The launch is carried entirely by curated sample images. Everything quantitative below comes from third parties.

arena.ai text-to-image (retrieved September 6, 2026)

RankModelEloVotesArena's licence label
#1gpt-image-2 (medium)1382ยฑ477,830Proprietary
#8seedream-5.0-pro1258ยฑ459,484Proprietary
#9qwen-image-3.0-pro1254ยฑ710,923Proprietary
#16qwen-image-2.0-pro-2026-06-221191ยฑ612,016Proprietary
#39qwen-image-2512โ€”โ€”Apache 2.0
#60qwen-imageโ€”โ€”Apache 2.0

Two things follow, both checkable at arena.ai/leaderboard/text-to-image:

  1. The generational gain is real: +63 Elo over 2.0-pro, and a jump from #16 to #9. That is a larger step than 2.0 made over the open checkpoints.
  2. "#1 among Chinese models" no longer holds. Chinese coverage at the August 5 rollout reported qwen-image-3.0-pro as fifth globally and first among Chinese models. A month later it is #9, and ByteDance's Seedream 5.0 Pro is one rank above it at #8. Arena rank claims decay in weeks; treat any quoted rank as a dated snapshot, including this one.

Artificial Analysis image arena (retrieved September 6, 2026)

A different organisation, a different Elo scale, and the two are not comparable with the arena.ai figures above. Artificial Analysis places Qwen-Image-3.0-Pro at #12 with an Elo of 1086 (5,551 samples) and Qwen-Image-3.0 at #14 with 1078 (5,177), against GPT Image 2 (high) at #1 with 1178.

Image editing

There is no Qwen-Image 3.0 entry on arena.ai's image-edit board as of September 6, 2026, even though the model is sold as an editor. The highest-placed Qwen editor there is still qwen-image-2.0-pro-2026-06-22 at #19 (1304ยฑ5, 50,824 votes), and the open qwen-image-edit sits at #31 with 1,981,112 votes โ€” roughly 39x the flagship's, and the most of any Apache 2.0 model on that board.

Limitations

  • Closed weights, again. No self-hosting, no fine-tuning on your own hardware, no offline or air-gapped inference, no LoRA ecosystem. Two generations in, this is policy rather than an exception.
  • No technical report and no model card. 2.0 at least produced a paper three months after launch. For 3.0 there is no architecture, no parameter count, no training description and no evaluation methodology โ€” you cannot reason about cost, latency or capability headroom from first principles.
  • No vendor benchmarks at all. Every number on this page was produced by someone other than Alibaba, on preference arenas rather than task benchmarks.
  • The 2K ceiling did not move. Total output pixels still cap at 2048x2048, the same as 2.0. Seedream 5.0's API reaches 4096x4096.
  • No pinnable snapshots. qwen-image-3.0-pro is an unversioned alias. The 2.0 series let you pin a dated build; here the model behind the ID can change without a name change, which makes your own evaluations expire silently.
  • Pro is rate-limited to 5 RPM on Alibaba's published limits, so batch workloads land on the cheaper standard tier or on a quota increase request.
  • 2K doubles the Pro price. The 1K/2K tier split means the resolution you want is a 2x cost decision on Pro, where the standard model charges one flat rate.
  • The 4,500-token limit is "recommended", not specified. Alibaba does not document what happens above it; the 2.0 series silently truncated over-long prompts.
  • Thin arena sample. 10,923 votes on text-to-image is a fraction of the 150,000+ behind the Gemini and GPT entries it is ranked against, so the placement will move.
  • Content policy constraints. As a model served primarily from Chinese-mainland infrastructure, generation is shaped by content rules that differ from those governing US-hosted models.

Pricing & Access

Alibaba bills image models per image, never per token, and 3.0 adds a 1K/2K resolution split and a charge for reference images that the 2.0 launch card did not have. The Beijing rate card is native CNY; Alibaba's international regions restate it in CNY at an applied exchange rate, which is why the Singapore column below is full of long decimals.

ModelBeijing (ๅ…ƒ/ๅผ )Singapore, as Alibaba quotes it (ๅ…ƒ/ๅผ )
qwen-image-3.0-pro, 1K outputยฅ0.25ยฅ0.299768
qwen-image-3.0-pro, 2K outputยฅ0.5ยฅ0.562065
qwen-image-3.0, 1K or 2K outputยฅ0.18ยฅ0.224826
Reference image input, either modelยฅ0.02ยฅ0.022483
qwen-image-2.0-pro outputยฅ0.5โ€”
qwen-image-2.0 outputยฅ0.18โ€”

In USD, the native quotes come from Qwen Cloud, the international surface Alibaba routes non-mainland developers to: $0.04 per 1K image and $0.075 per 2K image for Pro, $0.03 flat for the standard model, and $0.003 per input image. Those are the rates OpenRouter publishes for the Alibaba Cloud International provider, and they match Artificial Analysis's listing of $40 and $30 per 1,000 images. The round numbers are the native quotes; the long decimals in the table above are conversions.

A free quota is listed on the Chinese pricing page โ€” 10 images, input and output combined, valid for 90 days from activation. That is a much smaller allowance than the 100-image quota advertised during the 2.0 era, and Alibaba's regional pages have contradicted each other about free quotas before. Confirm it in your own console before relying on it.

Access routes: Alibaba Cloud Model Studio / Bailian and DashScope, Qwen Cloud, OpenRouter (Alibaba Cloud International), and Qwen Chat to try it interactively. There is no self-hosted path. Rates change โ€” check the pricing page before budgeting.

Ecosystem & Tools

Community & Resources

Frequently Asked Questions

Alibaba's Qwen team announced it on July 21, 2026, initially as an invite-only API preview; Alibaba Cloud's community repost of the announcement is dated July 22, 2026. General availability followed in early August 2026 โ€” Chinese tech press reported the full rollout on August 5, 2026, and OpenRouter lists the Pro model as added on August 5 with Alibaba Cloud International as the sole provider.
No. There is no Qwen-Image-3.0 repository under huggingface.co/Qwen โ€” the Hugging Face API lists seven Qwen image repositories on September 6, 2026, the newest of which is from May 2026 โ€” and ModelScope's API returns "record not found" for Qwen/Qwen-Image-3.0 while returning success for Qwen/Qwen-Image. That makes 3.0 the second consecutive closed Qwen image flagship, after Qwen-Image 2.0.
Alibaba does not publish one, so this page does not state one. There is no technical report, no model card and no architecture disclosure for 3.0 โ€” a break even from Qwen-Image 2.0, which at least shipped a paper three months after launch. Third-party sites that quote a parameter count for the 2.0 or 3.0 flagships are not citing Alibaba.
Model Studio's 3.0 API reference gives a recommended maximum of 4,500 tokens for the prompt. That is the headline generational change: the Qwen-Image 2.0 technical report described instructions "of up to 1K tokens" and the 2.0 API capped prompts at 1,300 tokens, itself already the highest ceiling in the Qwen image line.
Billing is per image and split by resolution tier. On Alibaba's Beijing rate card qwen-image-3.0-pro is ยฅ0.25 per 1K image and ยฅ0.5 per 2K image, qwen-image-3.0 is a flat ยฅ0.18, and reference images cost ยฅ0.02 each. Qwen Cloud quotes the same models natively in dollars at $0.04 (1K) and $0.075 (2K) for Pro, $0.03 for the standard tier, and $0.003 per input image.
Pro is the quality tier and the only one with a separate 2K price; the standard model charges one flat rate at both 1K and 2K. Alibaba's model pages also give them different throughput ceilings โ€” 5 requests per minute for Pro, 20 for the standard model โ€” and arena.ai lists only the Pro variant.
No. Model Studio documents the same ceiling as the 2.0 series: total output pixels between 512x512 and 2048x2048, with aspect ratios from 1:8 to 8:1. A generational bump did not raise the resolution ceiling, and Seedream 5.0's API still reaches 4096x4096.
On September 6, 2026, qwen-image-3.0-pro sits at #9 on the arena.ai text-to-image leaderboard with an Elo of 1254ยฑ7 from 10,923 votes, 63 points above qwen-image-2.0-pro-2026-06-22 at #16. It is not the top-placed Chinese model โ€” ByteDance's seedream-5.0-pro is one rank above it at #8.
Seven, all Apache 2.0 on Hugging Face: Qwen-Image, Qwen-Image-2512, Qwen-Image-Edit, Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-Layered and the Qwen-Image-Bench evaluation model. Qwen-Image-Edit-2509 leads at roughly 458,000 downloads in the 30 days to September 6, 2026. None of them is newer than May 2026.
They are separate model lines from separate teams. Qwen-Image comes from the Qwen team; Z-Image comes from Alibaba's Tongyi-MAI group. Tongyi-MAI/Z-Image-Turbo is Apache 2.0 and recorded roughly 677,000 Hugging Face downloads in the 30 days to September 6, 2026 โ€” more than any Qwen-Image checkpoint.

Explore More Models

Discover other AI models and compare their capabilities.