---
source: 'https://howaiworks.ai/models/vidu'
section: models
title: Vidu Q3 Pro
description: >-
  ShengShu's flagship Vidu Q3 Pro video model, announced January 30, 2026.
  Generates up to 16 seconds of 1080p video with native audio in a single pass.
tags:
  - Vidu
  - ShengShu Technology
  - Video Generation
  - AI Video
  - Reference-to-Video
  - Native Audio
  - Vidu S1
  - Chinese AI
  - Latest
category: Video Generation Models
developer: ShengShu Technology
developerWebsite: 'https://www.vidu.com/'
modelType: Video Generation Model
releaseDate: '2026-01-30'
lastUpdated: '2026-07-08'
license: Proprietary
officialWebsite: 'https://www.vidu.com/vidu-q3'
docsPage: 'https://platform.vidu.com/docs'
pricingPage: 'https://platform.vidu.com/docs/pricing'
---

# Vidu Q3 Pro

> ShengShu's flagship Vidu Q3 Pro video model, announced January 30, 2026. Generates up to 16 seconds of 1080p video with native audio in a single pass.

## Overview

Vidu Q3 Pro is the top tier of ShengShu Technology's Q3 video generation family, **announced on January 30, 2026**. ShengShu announced it from Singapore during Global Creativity Week, calling Vidu Q3 "the industry's first long-form AI video model to deliver native audio and video generation in a single output." The model produces up to 16 seconds of native 1080p [video generation](https://howaiworks.ai/glossary/video-generation) output with dialogue, sound effects, and music generated in the same pass.

ShengShu's pitch is consistency rather than raw fidelity. On **April 13, 2026** the company extended Q3 with reference-to-video, which lets a creator combine "subjects, environments, costumes, props, and visual styles" from uploaded references inside one generation. ShengShu states that "Vidu Q3 ranked No.1 in the first global Reference-to-Video leaderboard released by SuperCLUE," a Chinese [benchmark](https://howaiworks.ai/glossary/benchmark) organization; the leaderboard is described only in ShengShu's release and this page could not reproduce it. Reference-to-video runs on `viduq3-mix`, `viduq3-turbo`, and `viduq3` — not on `viduq3-pro`. That release also announced ShengShu's **RMB 2 billion Series B, led by Alibaba Cloud**.

The company frames all of this as a step toward a world model. Founder and chief scientist Dr. Zhu Jun: **"At its core, a world model gives AI a unified way to represent and predict the real world. Video plays a critical role in this, as it naturally captures time, space, motion, and causality."**

### Vidu S1 — the newest model, and not a successor

On **July 3, 2026**, five days before this page was last updated, ShengShu "unveiled Vidu S1, its next-generation video [foundation model](https://howaiworks.ai/glossary/foundation-models)" at the 2026 Global Digital Economy Conference. (The release is datelined Singapore but does not name the conference's host city.) S1 does something Q3 Pro cannot: continuous, real-time, voice-driven interaction with a generated character.

- **"The model delivers 540P (960x540) resolution at 25 FPS (up to 42 FPS)"**
- It uses an **autoregressive diffusion (AR + Diffusion)** architecture that "continuously predicts and generates subsequent video content based on previously generated frames, current voice instructions, and conversational context," rather than rendering a finished clip up front.
- **"Notably, the entire system runs on consumer-grade GPUs, significantly reducing the hardware requirements for real-time interactive video generation."**
- An interactive character is created from a single image, with no per-character modelling or training.
- [Inference](https://howaiworks.ai/glossary/inference) acceleration comes from TurboDiffusion, low-bit SageAttention, and sparse [attention](https://howaiworks.ai/glossary/attention-mechanism); serving is handled by ShengShu's own engine, **TurboServe**.

S1 is a different product line, not a Q3 Pro upgrade. Q3 Pro renders 1080p clips; S1 streams 540p interaction. ShengShu's S1 announcement does not mention Q3 at all, and S1 does not appear on the API model map. Try it at [vidu.com/vidu-stream](https://www.vidu.com/vidu-stream).

### A note on naming

**`viduq3.com` is not ShengShu's website.** It is a third-party reseller-style site whose own footer states it "is an independent AI tool website and is not affiliated with, endorsed by, or sponsored by any original model provider or trademark owner." Its credit packs ($19.90 to $99.90) are its own, not ShengShu's. The first-party surfaces are `vidu.com` (consumer) and `platform.vidu.com` (API).

**"Vidu Q3 Pro" is an API tier, not a separately launched model.** ShengShu launched "Vidu Q3" on January 30, 2026 and exposes it through four model IDs — `viduq3-pro`, `viduq3-turbo`, `viduq3-mix`, and `viduq3`. No separate launch date for the Pro tier is published, so this page dates the model to the Q3 launch.

## Capabilities

- **Native audio-video in one pass**: ShengShu's model map states that `viduq3-pro` "supports simultaneous audio and visual output." Dialogue, voiceover, sound effects, and music are generated with the picture rather than dubbed onto it afterward.
- **Reference-to-video**: 1 to 7 reference images condition a generation on specific subjects, costumes, props, environments, and styles. This is Vidu's oldest and still its clearest differentiator — it predates the Q3 line, and ShengShu marketed the Q2 version as a model "where anything can serve as a reference."
- **16-second clips**: `viduq3-pro` supports 1–16 seconds at 24fps, against a market where 5–10 seconds is common.
- **Cinematic effects**: the April 2026 release added six effect classes — particle systems, fluid simulation, dynamic motion, camera movement, transitions, and lighting.
- **Layered sound design**: five audio categories — ambient sound, motion-driven audio, atmospheric layers, foley effects, and emotion-driven cues.
- **Multilingual speech with lip sync**: the January release credits Q3 with "multilingual voice generation, precise lip synchronization"; the April release adds "multilingual dialogue." ShengShu publishes no list of supported languages.
- **Real-time interaction (via S1, not Q3)**: unlimited-duration, voice-driven character interaction at 540P.

## Technical Specifications

Model IDs and limits below are taken from ShengShu's own API model map and reference-to-video reference page. The two disagree: the model map's ViduQ3 table lists only `viduq3-pro` and `viduq3-turbo` and marks reference-to-video unsupported on both, while the reference-to-video endpoint and the pricing page both accept `viduq3-turbo`, `viduq3-mix`, and `viduq3` for it.

| Model ID | Capabilities | Resolutions | Duration | Frame rate |
|---|---|---|---|---|
| `viduq3-pro` | image-to-video, start-end-to-video, text-to-video | 540p / 720p / 1080p | 1–16s | 24fps |
| `viduq3-turbo` | image-to-video, start-end-to-video, text-to-video, reference-to-video | 540p / 720p / 1080p | 1–16s | 24fps |
| `viduq3-mix` | reference-to-video (1–7 reference images) | 720p / 1080p | 1–16s (default 5) | 24fps |
| `viduq3` | reference-to-video | 540p / 720p / 1080p | 3–16s (default 5) | 24fps |
| `viduq2-pro` | image-to-video, reference-to-video, start-end-to-video | 540p / 720p / 1080p | 1–10s | 24fps |
| `vidu2.0` | image-to-video, reference-to-video, start-end-to-video | 360p / 720p / 1080p | 4s / 8s | 32fps |

- **Aspect ratios**: the reference-to-video endpoint defaults to 16:9 and lists 16:9, 9:16, 3:4, 4:3, and 1:1 — but notes that 3:4 and 4:3 are supported only on the q2 model, so `viduq3-mix`, `viduq3`, and `viduq3-turbo` are limited to 16:9, 9:16, and 1:1.
- **License**: Proprietary. Closed weights.
- **Architecture (Q3)**: not published. ShengShu's public architecture lineage is **U-ViT** ("Universal Vision Transformer"), which Global Times reported "Vidu's research team" proposed in September 2022 and used for the original Vidu launched April 27, 2024. ShengShu has not confirmed that Q3 uses U-ViT.
- **Architecture (S1)**: autoregressive diffusion (AR + Diffusion), per ShengShu's announcement.
- **Parameters, training data, knowledge cutoff**: ShengShu publishes none of these, so this page does not state any.

## Use Cases

- **Character-consistent narrative video**: the reference-to-video path holds a cast, wardrobe, and set across shots — the workflow Vidu is actually built around.
- **Advertising and e-commerce creative**: ShengShu says commercial projects account for "over 70 percent of total output" on the platform.
- **Anime and comic-drama production**: ShengShu lists animation and comic drama among the industries adopting the Vidu API.
- **Short-form social video with sound**: 16 seconds of 1080p with music and foley in one generation, no audio post.
- **Real-time avatars, livestreaming, and game NPCs (S1)**: ShengShu names AI companions, virtual influencers, interactive livestreaming, game NPCs, customer service, and XR.
- **Cost-sensitive batch generation**: `viduq3-turbo` at 540p costs 7 credits/second ($0.035), under a third of `viduq3-pro` at 1080p (24 credits/second).

## Performance / Benchmarks

### arena.ai — Image-to-Video Arena

Verified directly against the arena.ai leaderboard (42 models, 1,350,288 votes, snapshot dated June 23, 2026). arena.ai is the former LMArena; its Elo scale is not comparable to Artificial Analysis's.

| Rank | Model | Score | Votes |
|---|---|---|---|
| 1 | `dreamina-seedance-2.0-720p` (ByteDance) | 1474±10 | 81,746 |
| 12 | **`vidu-q3-pro`** | **1361±8** | **36,677** |
| 27 | `vidu-q2-turbo` | 1243±17 | 2,506 |
| 33 | `vidu-q2-pro` | 1222±17 | 2,608 |

**No Vidu model appears on arena.ai's Text-to-Video Arena** (42 models, 533,418 votes, July 5, 2026). Vidu's public arena standing is image-conditioned only — consistent with a product built around references and stills rather than pure prompting.

### Vendor-reported placements

These are ShengShu's claims, made at launch, and this page could not independently confirm any of them:

| Claim | Source | Date |
|---|---|---|
| "Vidu Q3 ranked No.1 in China and No.2 globally" (Artificial Analysis) | ShengShu press release | Jan 30, 2026 |
| "At launch, Vidu Q3 ranked No.1 globally on the benchmark published by Artificial Analysis" | ShengShu press release | Apr 13, 2026 |
| "Vidu Q3 ranked No.1 in the first global Reference-to-Video leaderboard released by SuperCLUE" | ShengShu press release | Apr 13, 2026 |

Artificial Analysis publishes no dated archive of its video leaderboards, so neither launch-day placement can be checked against the record today. Note also that they contradict each other about the *same* moment: the January launch release said Q3 ranked "No.1 in China and No.2 globally," while the April release says that **"At launch"** Q3 ranked "No.1 globally." Treat both as vendor claims.

## Limitations

- **All ranking claims except the arena.ai figures are vendor-reported.** ShengShu's Artificial Analysis and SuperCLUE placements come from its own press releases, were true (if at all) only at launch, and could not be reproduced against a live leaderboard.
- **No model card.** No parameter count, no architecture paper for Q3, no training-data description, no knowledge cutoff. The U-ViT lineage is documented for the 2024 original, not for Q3.
- **Closed weights.** Nothing to self-host. No ShengShu repository exists on Hugging Face.
- **Absent from text-to-video evaluation.** Vidu has no entry on arena.ai's Text-to-Video Arena, so its prompt-only quality has no independent public measurement.
- **S1 is 540p.** Real-time interaction is bought at a resolution roughly one quarter the linear scale of Q3 Pro's 1080p, and ShengShu publishes no API pricing for it.
- **Consumer pricing is not machine-readable.** vidu.com/pricing loads its plan table client-side ("Loading pricing plans…"), so the subscription tiers cannot be verified from the page source. Third-party pages quoting Vidu subscription prices are second-hand.
- **Reference-to-video quality is unmeasured externally.** SuperCLUE's reference-to-video leaderboard is described only in ShengShu's own release; no independent replication is available.
- **Corporate domains are scattered.** As of July 8, 2026, `shengshu.com` 301-redirects to `www.shengshu.com`, which 302-redirects to `www.genspi.com` — a live ShengShu corporate site listing Vidu S1, Vidu Q3, and Motubrain. Meanwhile `shengshu-ai.com` presents an expired TLS certificate. ShengShu has published no explanation, so this page cites `vidu.com` as the developer website and draws no conclusion about a rebrand.

## Pricing & Access

### API pricing

ShengShu's [pricing page](https://platform.vidu.com/docs/pricing) states: **"Credits are available for $0.005 each, with applicable sales tax based on your location."** Video is billed per second of output.

| Model | 1080p | 720p | 540p |
|---|---|---|---|
| `viduq3-pro` | 24 credits/s ($0.12) | 20 credits/s ($0.10) | 9 credits/s ($0.045) |
| `viduq3-pro` off-peak | 12 credits/s ($0.06) | 10 credits/s ($0.05) | 5 credits/s ($0.025) |
| `viduq3-turbo` | 13 credits/s ($0.065) | 11 credits/s ($0.055) | 7 credits/s ($0.035) |
| `viduq3-turbo` off-peak | 7 credits/s ($0.035) | 6 credits/s ($0.03) | 4 credits/s ($0.02) |
| `viduq3-pro-fast` (image-to-video only) | 25 credits/s ($0.125) | 20 credits/s ($0.10) | — |

Reference-to-video on `viduq3-mix` is priced separately: **29 credits/second at 1080p ($0.145) and 24 credits/second at 720p ($0.12)**. Off-peak discounting is *not* supported for `viduq3-mix`.

Off-peak mode is roughly half price on the models that support it. A 16-second 1080p `viduq3-pro` generation therefore costs about **$1.92** at standard rates, or **$0.96** off-peak.

For scale: [Sora 2](https://howaiworks.ai/models/sora) is $0.10/second at 720p, so Vidu's 720p tier is exactly at parity — but Vidu allows 16-second generations where `sora-2` does not.

### Consumer and enterprise access

- **Vidu App and vidu.com** — the consumer surface. Plan prices are rendered client-side and are not stated here.
- **Vidu Agent, Vidu Claw** — ShengShu's production tooling; the April 2026 release says that, built on the Q3 model foundation, "Vidu has been fully integrated across its product ecosystem, including Vidu Agent, Vidu Claw, and the Vidu App."
- **Alibaba Cloud Model Studio** — Q3 is available through ShengShu's Series B lead investor's model platform.
- **Vidu Stream** — the S1 surface, at [vidu.com/vidu-stream](https://www.vidu.com/vidu-stream); the S1 release gives `platform.vidu.com/live/landing` as the developer API entry point.

ShengShu reports Vidu reaching "users in more than 200 countries and regions, serving over 40 million creators and more than 10,000 developers and enterprise customers," with "more than 500 million videos… generated on the platform." These are company figures, unaudited.

## Ecosystem & Tools

- **[Vidu API platform](https://platform.vidu.com/)** — MaaS access, credit billing, off-peak mode
- **[API model map](https://platform.vidu.com/docs/model-map)** — ShengShu's model ID table; incomplete, as it omits `viduq3-mix`, `viduq3`, and `viduq3-pro-fast`
- **[Reference-to-video reference](https://platform.vidu.com/docs/reference-to-video)** — `viduq3-mix` parameters and the 1–7 reference-image limit
- **[Vidu Stream](https://www.vidu.com/vidu-stream)** — the Vidu S1 real-time interactive surface
- **Alibaba Cloud Model Studio** — third-party hosted access to Q3
- **Vidu Agent 1.0 / Reference Hub 2.0** — storyboard editing, narration removal, and reference management, announced January 30, 2026
- **Motubrain** — ShengShu's "world-action model for the physical world," listed alongside Vidu on the company site; its relationship to the Vidu architecture is not published ([embodied AI](https://howaiworks.ai/glossary/embodied-ai))

## Community & Resources

- [ShengShu Technology Unveils Vidu S1](https://www.prnewswire.com/news-releases/shengshu-technology-unveils-vidu-s1-bringing-real-time-interactive-generation-to-ai-video-302817626.html) — the July 3, 2026 announcement
- [ShengShu Launches Vidu Q3 Reference-to-Video](https://www.prnewswire.com/news-releases/shengshu-launches-vidu-q3-reference-to-video-with-expanded-visual-and-audio-capabilities-302740489.html) — April 13, 2026, with the Zhu Jun world-model quote and the Series B
- [Vidu Showcases "China Speed" at Global Creativity Week](https://www.prnewswire.com/news-releases/vidu-showcases-china-speed-in-advancing-ai-video-into-production-at-global-creativity-week-302675040.html) — the January 30, 2026 Vidu Q3 launch
- [Chinese AI Startup ShengShu Bags USD293 Million](https://www.yicaiglobal.com/news/chinese-ai-startup-shengshu-bags-usd293-million-in-latest-fundraiser-led-by-alibaba-cloud) — Yicai Global on the Alibaba Cloud-led Series B
- [arena.ai Image-to-Video Arena](https://arena.ai/leaderboard/image-to-video) — where `vidu-q3-pro` actually sits
- [Chinese team unveils first text-to-video AI model on par with Sora](https://www.globaltimes.cn/page/202404/1311367.shtml) — Global Times on the original Vidu and U-ViT, April 2024
- Compare with [Sora 2](https://howaiworks.ai/models/sora), [Kling 3.0](https://howaiworks.ai/models/kling), [Seedance 2.0](https://howaiworks.ai/models/seedance), [Hailuo 2.3](https://howaiworks.ai/models/hailuo), [Wan 2.2](https://howaiworks.ai/models/wan), and [HappyHorse 1.1](https://howaiworks.ai/models/happyhorse)

## Frequently Asked Questions

### When was Vidu Q3 released?

January 30, 2026. ShengShu Technology announced it from Singapore during Global Creativity Week, describing Vidu Q3 as "the industry's first long-form AI video model to deliver native audio and video generation in a single output." A reference-to-video capability was added to the Q3 line on April 13, 2026.

### Is Vidu S1 a newer version of Vidu Q3 Pro?

It is newer — announced July 3, 2026 — but it is not a replacement. ShengShu calls S1 "its next-generation video foundation model," yet S1 runs at 540P (960x540) for real-time interactive sessions, while Q3 Pro produces up to 16 seconds of 1080p clip output. They are separate product lines: the Q series generates clips, the S series streams interaction.

### Is Vidu open source?

No. Vidu is proprietary and closed weights. There are no official ShengShu or Vidu model repositories on Hugging Face, and ShengShu publishes no model card, parameter count, or training-data description for the Q3 series. Access is through the Vidu apps or the Vidu API.

### How much does Vidu Q3 Pro cost?

The Vidu API is credit-based. ShengShu's pricing page states that "credits are available for $0.005 each." `viduq3-pro` costs 24 credits per second at 1080p ($0.12/second), 20 at 720p ($0.10/second), and 9 at 540p ($0.045/second). An off-peak mode roughly halves those rates.

### What is reference-to-video and why is it Vidu's differentiator?

Reference-to-video conditions generation on uploaded reference images rather than on a prompt alone, so specific characters, costumes, props, and environments stay consistent across a clip. Vidu's API accepts 1 to 7 reference images on `viduq3-mix`. Note that `viduq3-pro` itself does not expose reference-to-video: the endpoint accepts `viduq3-mix`, `viduq3-turbo`, `viduq3`, and older IDs. ShengShu says Vidu Q3 "ranked No.1 in the first global Reference-to-Video leaderboard released by SuperCLUE" — a vendor claim this page could not reproduce.

### Where does Vidu rank on arena.ai?

`vidu-q3-pro` sits at rank 12 of 42 on the arena.ai Image-to-Video Arena with a score of 1361±8 across 36,677 votes. No Vidu model appears on arena.ai's Text-to-Video Arena at all.

### Is viduq3.com the official Vidu website?

No. viduq3.com is a third-party site whose own footer states it "is an independent AI tool website and is not affiliated with, endorsed by, or sponsored by any original model provider or trademark owner." Its $19.90–$99.90 credit packs are not ShengShu's prices. The official surfaces are vidu.com and platform.vidu.com.

### Does Vidu Q3 generate audio?

Yes, natively. ShengShu's API model map says `viduq3-pro` "supports simultaneous audio and visual output." The April 2026 reference-to-video release added five categories of sound: ambient sound, motion-driven audio, atmospheric layers, foley effects, and emotion-driven cues.

### Who builds Vidu, and who funds it?

ShengShu Technology, founded in 2023, with a core team from Tsinghua University's Institute for Artificial Intelligence. Zhu Jun is founder and chief scientist; Yihang Luo is CEO. Yicai Global reported on April 10, 2026 that ShengShu had "secured nearly CNY2 billion (USD293 million)" in a Series B led by Alibaba Cloud; ShengShu's own release states the round was "RMB 2 billion."

### Is there a Vidu Q4 or Vidu 3.0?

Neither appears anywhere in ShengShu's API documentation. The newest Q-series model IDs are `viduq3-pro` and `viduq3-turbo` (the only two rows in the API model map's ViduQ3 table) plus `viduq3-mix` and `viduq3`, which appear on the reference-to-video endpoint and the pricing page. The newest model ShengShu has announced is Vidu S1, which is not a Q-series successor.

## Related

### Related models

- [Sora 2 (Discontinued)](https://howaiworks.ai/models/sora)
- [Kling 3.0](https://howaiworks.ai/models/kling)
- [Seedance 2.5](https://howaiworks.ai/models/seedance)
- [MiniMax H3 (Hailuo 3.0)](https://howaiworks.ai/models/hailuo)
- [Wan 3.0](https://howaiworks.ai/models/wan)
- [HappyHorse 1.1](https://howaiworks.ai/models/happyhorse)

---

Source: https://howaiworks.ai/models/vidu — HowAIWorks.ai
