Wan 3.0

Alibaba's video family. Wan 3.0 (August 2026) generates 30-second 1080p clips with audio, API-only. Wan 2.2 is still the last Apache 2.0 release.

Updated

Released
Aug 1, 2026
Type
Video Generation Model
License
Proprietary (Wan 3.0); Apache 2.0 (Wan 2.2 open weights)
On this page

Overview

Wan is Alibaba's video generation family, built by Tongyi Lab and marketed in China as ι€šδΉ‰δΈ‡η›Έ (Tongyi Wanxiang). It is two products wearing one name, and telling them apart is the entire job of this page.

Wan 3.0 is the current flagship, and it is closed. Alibaba put it into public beta in early August 2026 and published its own write-up on August 13, 2026, describing "up to 30 seconds of video in a single pass" at 480P, 720P or 1080P, with audio generated alongside the picture. Its distinguishing input is breadth: alongside text, images, audio and video it accepts documents β€” doc, xls, ppt, pdf, txt, md β€” and public webpages, and builds a video from what it reads. Model Studio serves it as wan3.0-video and wan3.0-video-prime. There are no weights.

Wan 2.2 is the current open model, and it is from July 2025. Its headline architectural claim is that "Wan2.2 introduces a Mixture-of-Experts (MoE) architecture into video diffusion models." The two large variants carry 27B total parameters with 14B active per denoising step, split into "a high-noise expert for the early stages, focusing on overall layout; and a low-noise expert for the later stages, refining video details." Alongside them Alibaba shipped a 5B dense model, TI2V-5B, built on a Wan2.2-VAE (autoencoder) with a 4Γ—16Γ—16 compression ratio, which "can generate a 5-second 720P video in under 9 minutes on a single consumer-grade GPU."

That 5B model is why Wan matters to most of the people reading this. It made competent local video generation a consumer-hardware activity, and the community built accordingly: Comfy-Org/Wan_2.2_ComfyUI_Repackaged alone records roughly 4.85 million Hugging Face downloads in the last 30 days, more than a year after release. The Wan2.2 GitHub repository has 17.4k stars.

The version ladder, and where the weights stop

VersionDateWeights
Wan 2.1February 2025Apache 2.0
Wan 2.2July 28, 2025Apache 2.0
Wan 2.5-PreviewSeptember 24, 2025None
Wan 2.6December 16, 2025None
Wan 2.7April 2026None
Wan 3.0August 2026None

Four consecutive closed generations. Alibaba's announcements are consistent about this by omission: the Wan2.6 press release routes users to "Model Studioβ€”Alibaba Cloud's AI development platformβ€”and Wan's official website," the Wan2.7-Video post says both models are "now available on Alibaba Cloud's Model Studio and the official Wan website," and the Wan3.0 post says it is "now available on Alibaba Cloud Model Studio." None of the three mentions open source, weights or a licence. A Hugging Face hub-wide search for Wan2.7 returns zero models; the same is true of Wan3. GitHub issue #184 ("Is WAN 2.5 going to be open source?") and #181 were both opened on September 24, 2025 and both remain open with no maintainer answer.

The one qualification: Tongyi Lab has not stopped open-sourcing entirely. It has stopped open-sourcing the frontier line. In July and August 2026 it published two new Apache 2.0 models built on the Wan 2.2 base β€” Wan-Animate-2 and Wan-Dancer-14B β€” which is a real signal about where the open branch is allowed to go: derivative and specialised, not frontier.

Treat any benchmark table that says only "Wan" as unresolved until you know which one it means.

Capabilities

Wan 3.0 (hosted, no weights)

  • One model, all tasks: wan3.0-video is an all-in-one reference model covering text-to-video, image-to-video (first frame and first/last frame) and reference-based generation, rather than a suite of separate endpoints.
  • 30-second clips: 2–30 seconds at 480P, 720P or 1080P, 30fps. Setting duration: -1 hands length selection to the model.
  • Native audio: on by default; audio: false disables it, and Alibaba states this does not change the price.
  • Document and webpage input: one file (up to 100MB, 50 pages) or one public URL per generation, read before generating.
  • wan3.0-video-prime: the same capabilities tuned for speed.

Wan 2.7 (hosted, no weights — superseded April→August 2026)

Four separate models, which is the structure Wan 3.0 collapsed:

  • wan2.7-t2v β€” text-to-video, 2–15s, 720P/1080P, audio sync and multi-shot narrative.
  • wan2.7-i2v β€” image-to-video, 2–15s, with first-frame, first-and-last-frame and video-continuation tasks.
  • wan2.7-r2v β€” reference-to-video, 2–10s, up to five mixed image/video/audio references. Alibaba's reference documentation notes that only Wan 2.7 supports voice reference: a reference_voice audio URL attached to a reference image or video carries that speaker's timbre into the output.
  • wan2.7-videoedit β€” instruction-based editing and style transfer, 2–10s, no audio.

Wan 2.2 (Apache 2.0, downloadable)

  • Text-to-video: Wan2.2-T2V-A14B generates 5-second clips at 480P and 720P.
  • Image-to-video: Wan2.2-I2V-A14B animates a still image at 480P and 720P.
  • Unified text-and-image-to-video on consumer hardware: Wan2.2-TI2V-5B covers both tasks in one 5B model at 720P/24fps on a 24GB GPU.
  • Speech-to-video: Wan2.2-S2V-14B is "an audio-driven cinematic video generation model" β€” it animates a portrait to a supplied audio track.
  • Character animation: Wan2.2-Animate-14B, and its August 2026 successor Wan2.2-Animate-2-14B, "directly consumes driving videos in a redesigned Diffusion Transformer."
  • Music-to-dance: Wan-Dancer-14B, "a hierarchical framework for minute-scale coherent music-to-dance generation."

Technical Specifications

Hosted models (Alibaba Cloud Model Studio):

Model IDTaskResolutionDurationAudio
wan3.0-videoAll-in-one480P / 720P / 1080P2–30sYes
wan3.0-video-primeAll-in-one, fast480P / 720P / 1080P2–30sYes
wan2.7-t2vText-to-video720P / 1080P2–15sYes
wan2.7-i2vImage-to-video720P / 1080P2–15sYes
wan2.7-r2vReference-to-video720P / 1080P2–10sYes
wan2.7-videoeditInstruction editing720P / 1080P2–10sNo
wan2.2-t2v-plusText-to-video480P / 1080P5sNo
wan2.2-i2v-plusImage-to-video480P / 1080P5sNo

Alibaba publishes no parameter count, architecture description or technical report for Wan 2.6, 2.7 or 3.0, so this page states none. Everything below is Wan 2.2, where the numbers are documented.

  • Open checkpoints (Apache 2.0): Wan-AI/Wan2.2-T2V-A14B, Wan-AI/Wan2.2-I2V-A14B, Wan-AI/Wan2.2-TI2V-5B, Wan-AI/Wan2.2-S2V-14B, Wan-AI/Wan2.2-Animate-14B, Wan-AI/Wan2.2-Animate-2-14B, Wan-AI/Wan-Dancer-14B β€” plus -Diffusers variants of most
  • Architecture: MoE diffusion transformer; two-expert denoising split by signal-to-noise-ratio threshold
  • Parameters: 27B total, 14B active per step (A14B models); 5B dense (TI2V-5B)
  • VAE: Wan2.2-VAE, 4Γ—16Γ—16 compression; TI2V-5B reaches 4Γ—32Γ—32 with an added patchification layer
  • VRAM: "at least 80GB VRAM" for single-GPU A14B inference; "at least 24GB VRAM (e.g, RTX 4090 GPU)" for TI2V-5B. Wan-Animate-2 is documented on 8Γ— A800 for 720P and 2Γ— A800 for 480P
  • Training data scale: "+65.6% more images and +83.2% more videos" than Wan 2.1 (relative; Alibaba publishes no absolute counts)
  • Technical report: arXiv:2503.20314, Wan: Open and Advanced Large-Scale Video Generative Models, which covers the 1.3B and 14B Wan 2.1 generation

Use Cases

  • Local video generation: still the dominant reason to be on this page. TI2V-5B on a 24GB consumer GPU, or a quantized A14B via GGUF, with no API dependency and no per-second billing β€” accepting a July 2025 quality ceiling.
  • ComfyUI production pipelines: node-based workflows chaining Wan 2.2 with upscalers, interpolators and LoRAs.
  • LoRA fine-tuning and character consistency: hundreds of Hugging Face repositories declare Wan2.2-I2V-A14B as their base model, most of them LoRAs.
  • Talking-head, avatar and dance video: Wan2.2-S2V-14B for audio-driven portrait animation, Wan2.2-Animate-2-14B for motion transfer and character replacement, Wan-Dancer-14B for music-synchronised dance β€” all Apache 2.0 and all recent.
  • Document-to-video via API: Wan 3.0's distinctive job. Turning a deck, a spreadsheet or a product page into a narrated 30-second clip is not something the open models can do at all.
  • Character-and-voice continuity via API: wan2.7-r2v's reference_voice, when a series needs the same face and voice across shots.
  • Research baselines: an Apache 2.0 MoE video diffusion model with a published technical report remains a rare, citable starting point.

Performance / Benchmarks

Alibaba's own claim, from the Wan2.2 model card, is "TOP performance among all open-sourced and closed-sourced models" on "our new benchmark Wan-Bench 2.0." Wan-Bench 2.0 is Alibaba's own benchmark, the claim dates from July 2025, and it has not aged.

Independent arenas tell a consistent story: the closed models are genuinely at the frontier, and the open one is not on the board with them.

Artificial Analysis Video Arena (blind pairwise votes, Elo; figures read September 2026). Its boards split into "With Audio" and "No Audio" views, and the two scales must not be cross-compared:

ModelText-to-video (with audio)Image-to-video (with audio)
Wan 3.0#2 β€” Elo 1239#5 β€” Elo 1175
Wan 2.7#11 β€” Elo 1106#12 β€” Elo 1083
Wan 2.6#25 β€” Elo 1028#33 β€” Elo 895

On the No Audio views the same source places Wan 3.0 first on text-to-video (Elo 1333) and second on image-to-video (Elo 1358). The open Wan 2.2 generates silent video and does not appear on the With Audio boards at all.

arena.ai runs a separate Elo scale, again not comparable with the above. Its own announcement of the Wan 3.0 listing placed it #3 on the Image-to-Video Arena at 1481 points with a 57% win rate, "a significant +53 pt improvement from Wan 2.7 at #10 with 1428 pts," 16 points behind first-place MiniMax-H3.

Two things follow. Wan 3.0 is a real frontier contender rather than a version bump β€” top-three on both major arenas, in a field led by Google, MiniMax and ByteDance. And every strong Wan result belongs to a model you cannot download: the open Wan 2.2 competes only on the No Audio boards, three generations behind its own closed siblings.

Limitations

  • The open model is four generations behind. Wan 2.2 dates from July 2025. Wan 2.5, 2.6, 2.7 and 3.0 have all shipped since, none with weights. If you need frontier quality you are using an API, and if you need weights you are using a year-old model.
  • No consumer-hardware path above 2.2. This is the single most important thing on the page for local users: there is no 5B Wan 3.0, no quantized Wan 2.7, no RTX 4090 route to either. The 24GB path is Wan2.2-TI2V-5B or a community GGUF, and nothing newer.
  • No audio generation in the open weights. Wan 2.2 T2V and I2V output silent video. S2V consumes audio, it does not synthesise it.
  • Serious hardware even for the open variants. Official single-GPU inference for the A14B models wants "at least 80GB VRAM"; Wan-Animate-2 is documented on multi-GPU A800 configurations.
  • Short clips locally. 5 seconds at 480P/720P, against 30 seconds at 1080P on the hosted 3.0.
  • wan2.7-videoedit has no audio, and both wan2.7-r2v and wan2.7-videoedit cap at 10 seconds rather than 15 β€” a detail that catches people who read only the headline number.
  • Thinking Mode is narrower than the marketing. The documented thinking_mode parameter belongs to the Wan-Image 2.7 models, not the video ones; the Wan3.0 video API exposes prompt_extend instead.
  • Nothing is published about the closed models' internals. No parameter counts, no architecture, no technical report for 2.6, 2.7 or 3.0.
  • The open-source commitment is not a commitment. Two GitHub issues asking whether Wan 2.5+ would be opened have sat unanswered since September 2025. Plan as though 2.2 is the last open frontier release, because for four generations it has been.
  • Apache 2.0 covers the weights only. Training data and Wan-Bench 2.0's contents are not published.

Pricing & Access

Self-hosting (Wan 2.2 and its Apache 2.0 descendants)

Free, with no revenue cap and no field-of-use restriction. Alibaba adds: "We claim no rights over the your generated contents." Weights are at Wan-AI on Hugging Face and ModelScope; inference code at Wan-Video/Wan2.2.

Hosted (all generations)

Alibaba Cloud bills video output by duration β€” "Cost = Video unit price Γ— Video duration (seconds)" β€” and prices by resolution. Note that for Wan 2.7 the billed duration is capped: min(10, rounded duration).

ModelResolutionFirst-party price
wan3.0-video480P$0.05 / second
wan3.0-video720P$0.10 / second
wan3.0-video1080P$0.20 / second
wan2.7-t2v / wan2.7-i2v720P$0.10 / second
wan2.7-t2v / wan2.7-i2v1080P$0.15 / second
wan2.2-t2v-plus / wan2.2-i2v-plus480P$0.02 / second
wan2.2-t2v-plus / wan2.2-i2v-plus1080P$0.10 / second

A full 30-second 1080p Wan 3.0 clip is therefore about $6.00. Artificial Analysis, quoting its own collected rates, lists Wan 3.0 at $12.00 per minute of output and Wan 2.7 at $9.00 per minute β€” consistent with the 1080P rows above.

Consumer access to the closed models runs through wan.video and, per Alibaba, the Qwen App.

Ecosystem & Tools

  • ComfyUI β€” official native workflows for TI2V-5B, T2V-A14B and I2V-A14B; the 5B "should fit well on 8GB vram with the ComfyUI native offloading". See also the ComfyUI tool page
  • Comfy-Org/Wan_2.2_ComfyUI_Repackaged β€” ~4.85M downloads in 30 days, the most-used Wan artifact on Hugging Face by a wide margin
  • GGUF quantizations β€” QuantStack/Wan2.2-T2V-A14B-GGUF (~526k/30d) and QuantStack/Wan2.2-I2V-A14B-GGUF (~224k/30d)
  • Diffusers β€” first-party -Diffusers checkpoints for the Wan 2.2 line, including Wan2.2-Animate-2-14B and a distilled variant
  • Wan-Video/Wan-skills β€” Apache 2.0 agent skills that call the closed Wan APIs; requires an Alibaba Cloud account and a Model Studio API key
  • Wan2.1 tool page β€” the earlier open generation, still in use for its 1.3B variant
  • Hosted inference β€” Alibaba Cloud Model Studio and Qwen Cloud for the full family; Together AI and other resellers for Wan 2.2

Community & Resources

Frequently Asked Questions

Wan 3.0, launched in August 2026. Alibaba's own announcement post is dated August 13, 2026 and describes "up to 30 seconds of video in a single pass" from text, images, audio, video or documents. Model Studio serves it as wan3.0-video and the speed-optimised wan3.0-video-prime. It supersedes Wan 2.7 from April 2026.
No. Wan 3.0 is API-only. Alibaba's announcement says it is "now available on Alibaba Cloud Model Studio" and mentions no weights, no repository and no licence. There is no Wan3 repository in the Wan-Video GitHub organization and no Wan-AI/Wan3* model on Hugging Face.
No, and this is the most widely repeated falsehood about the family. A Hugging Face hub-wide search for "Wan2.7" returns zero models, there is no Wan2.7 repository in the Wan-Video GitHub organization, and Alibaba's own Wan2.7-Video post says only that the models are "now available on Alibaba Cloud's Model Studio and the official Wan website." No open-source claim, no weights, no licence.
Only the Wan 2.1 and Wan 2.2 generations. The Wan-AI Hugging Face organization publishes nothing newer, and the Wan-Video GitHub organization's model repositories are Wan2.1, Wan2.2, Wan-Animate-2 and Wan-Dancer. Wan 2.5, 2.6, 2.7 and 3.0 have never shipped checkpoints.
A closed release announced on December 16, 2025 that added a reference-to-video model to the four existing ones. Alibaba's pitch was that wan2.6-r2v lets you "upload a character reference video with both appearance and voice" and keep "the distinctive look and sound of the original reference" in new scenes. Model Studio still serves wan2.6-t2v, wan2.6-i2v, wan2.6-i2v-flash, wan2.6-r2v and wan2.6-r2v-flash.
No. Neither has published weights, so there is no local path at any VRAM budget. The consumer-hardware story stops at Wan 2.2: Wan2.2-TI2V-5B runs on "a GPU with at least 24GB VRAM (e.g, RTX 4090 GPU)" and ComfyUI reports the 5B "should fit well on 8GB vram" with native offloading. If you came here for a local model, Wan 2.2 is the answer and it is thirteen months old.
The hosted models do; the downloadable ones do not. wan3.0-video generates audio by default and exposes an audio boolean to turn it off, and Alibaba's Wan2.7 text-to-video reference notes that without an input.audio_url "the model generates background music or sound effects that match the video content." The open Wan 2.2 T2V and I2V models output silent video; Wan2.2-S2V-14B consumes an audio track rather than synthesising one.
Self-hosting Wan 2.2 is free under Apache 2.0. Hosted video is billed per second of output and priced by resolution. Alibaba's Wan3.0 announcement lists $0.05/second at 480P, $0.10/second at 720P and $0.20/second at 1080P β€” about $6.00 for a full 30-second 1080p clip. Alibaba's rate card for Wan 2.7 lists $0.10/second at 720P and $0.15/second at 1080P.
A planning pass before generation, and it is narrower than the marketing suggests. Alibaba's Wan-Image 2.7 API reference documents a thinking_mode boolean, default true, supported only by wan2.7-image and wan2.7-image-pro, where it "enhances its reasoning capabilities to improve image quality" at the cost of generation time. The Wan3.0 video API documents prompt_extend (default true) instead, and the Wan 2.7 video API references document neither.
Yes, but only on the Wan 2.2 base. Wan-Animate-2 shipped Apache 2.0 weights on August 7, 2026 for character animation from a driving video, and Wan-Dancer-14B shipped on July 13, 2026 for "minute-scale coherent music-to-dance generation." Both are auxiliary models built on the 2.2 generation. The frontier text-to-video line has been closed since 2.2.

Explore More Models

Discover other AI models and compare their capabilities.