Alibaba Tongyi Lab has launched a public beta of Wan 3.0, the latest iteration of its video generation model family. Available on Alibaba Cloud Model Studio and Qwen Cloud under the model identifier wan3.0-video, the model produces up to 30 seconds of continuous video in a single pass at resolutions up to 1080p.
Unlike predecessor models such as Wan 2.7, which capped single-pass output at 15 seconds, Wan 3.0 consolidates video synthesis into a unified architecture and expands supported input modalities beyond prompt text and still images.

Omni-Reference Multimodal Inputs
The primary architectural addition in Wan 3.0 is support for structured document ingestion alongside traditional media references. The model accepts:
- Office documents, including PDF files, PowerPoint slide decks, and spreadsheets.
- Web pages, URL references, and raw HTML structures.
- Audio tracks and spoken-word voice clips for synchronized lip movement.
- Reference video clips, character images, and style templates.
By ingesting slide decks or structured reports directly, Wan 3.0 extracts sequential semantic information to generate explanatory or promotional video sequences without requiring manual prompt decomposition.
Visual Continuity and Resolution Tiers
To mitigate common temporal artifacts such as character drift, flickering, and background deformation across extended generation horizons, Wan 3.0 implements enhanced cross-frame attention mechanisms. Alibaba reports improved stability across facial micro-expressions, user interface elements, and text rendering in synthesized scenes.
Generation outputs are available in three resolution profiles:
- 480p standard definition for rapid prototyping.
- 720p high definition for general web playback.
- 1080p full high definition for final production delivery.
Alibaba has not released open weights for Wan 3.0, restricting access to API endpoints and cloud hosting services. Previous open-weight models in the series, such as Wan 2.1, remain available on open model repositories.
Enterprise and Robotics Applications
Beyond creative content production and marketing workflows, Alibaba is positioning Wan 3.0 as a simulation engine. The model is capable of generating synthetic video data to train autonomous driving perception models and humanoid robotics vision systems, simulating diverse lighting conditions, dynamic obstacles, and complex physical interactions.
Access is currently open in public beta through Alibaba Cloud Model Studio and Qwen Cloud, with commercial API pricing structured on a per-second video generation metric.



