Alibaba Launches Wan 3.0 AI Video Model with Native 30-Second Generation and Document Inputs

Alibaba Tongyi Lab has launched a public beta of Wan 3.0, the latest iteration of its video generation model family. Available on Alibaba Cloud Model Studio and Qwen Cloud under the model identifier wan3.0-video, the model produces up to 30 seconds of continuous video in a single pass at resolutions up to 1080p. Unlike predecessor models such as Wan 2.7, which capped single-pass output at 15 seconds, Wan 3.0 consolidates video synthesis into a unified architecture and expands supported input mo

2 min
Alibaba Launches Wan 3.0 AI Video Model with Native 30-Second Generation and Document Inputs

Alibaba Tongyi Lab has launched a public beta of Wan 3.0, the latest iteration of its video generation model family. Available on Alibaba Cloud Model Studio and Qwen Cloud under the model identifier wan3.0-video, the model produces up to 30 seconds of continuous video in a single pass at resolutions up to 1080p.

Unlike predecessor models such as Wan 2.7, which capped single-pass output at 15 seconds, Wan 3.0 consolidates video synthesis into a unified architecture and expands supported input modalities beyond prompt text and still images.

Multimodal Video Synthesis Architecture

Omni-Reference Multimodal Inputs

The primary architectural addition in Wan 3.0 is support for structured document ingestion alongside traditional media references. The model accepts:

  • Office documents, including PDF files, PowerPoint slide decks, and spreadsheets.
  • Web pages, URL references, and raw HTML structures.
  • Audio tracks and spoken-word voice clips for synchronized lip movement.
  • Reference video clips, character images, and style templates.

By ingesting slide decks or structured reports directly, Wan 3.0 extracts sequential semantic information to generate explanatory or promotional video sequences without requiring manual prompt decomposition.

Visual Continuity and Resolution Tiers

To mitigate common temporal artifacts such as character drift, flickering, and background deformation across extended generation horizons, Wan 3.0 implements enhanced cross-frame attention mechanisms. Alibaba reports improved stability across facial micro-expressions, user interface elements, and text rendering in synthesized scenes.

Generation outputs are available in three resolution profiles:

  • 480p standard definition for rapid prototyping.
  • 720p high definition for general web playback.
  • 1080p full high definition for final production delivery.

Alibaba has not released open weights for Wan 3.0, restricting access to API endpoints and cloud hosting services. Previous open-weight models in the series, such as Wan 2.1, remain available on open model repositories.

Enterprise and Robotics Applications

Beyond creative content production and marketing workflows, Alibaba is positioning Wan 3.0 as a simulation engine. The model is capable of generating synthetic video data to train autonomous driving perception models and humanoid robotics vision systems, simulating diverse lighting conditions, dynamic obstacles, and complex physical interactions.

Access is currently open in public beta through Alibaba Cloud Model Studio and Qwen Cloud, with commercial API pricing structured on a per-second video generation metric.

Sources

Written by

More to read

  • Hierarchical Tree-Organized Retrieval (RAPTOR) in Production RAG: Recursive Summarization, Gaussian Mixture Clustering, and Cross-Scale Querying

    Standard retrieval-augmented generation (RAG) architectures operate on flat document chunks. Corpora are split into fixed token windows (typically 256 to 1024 tokens), mapped into vector space via dense embedding models, and queried through approximate nearest neighbor (ANN) search. While this setup efficiently resolves localized factual lookups ("What is the termination clause in contract X?"), it systematically fails on thematic synthesis, cross-document comparison, and high-level aggregation

    1 min
  • Data Pruning and Core-Set Selection in Deep Learning: How EL2N, GraNd, and Memorization Dynamics Break Power-Law Scaling

    Modern foundation models are trained on tens of trillions of tokens, requiring millions of GPU hours. Yet empirical analyses consistently reveal that massive portions of web-crawled corpora and large-scale vision datasets are either highly redundant, uninformative, or dominated by unlearnable noise. Standard empirical scaling laws (such as those formulated by Kaplan et al. and Chinchilla) model generalization error as a power-law function of total training samples ($L(N) \propto N^{-\alpha}$). H

    1 min
  • XPeng Robotics Raises Over 00M at .3B Valuation to Scale Humanoid Robot Production

    Chinese electric vehicle manufacturer XPeng has announced that its robotics affiliate raised over $900 million in its first major institutional financing round. The investment values the robotics business at more than $6.3 billion post-money, representing one of the largest single private capital raises in the embodied AI sector to date. The round was led by IDG Capital and Gaorong Ventures, with participation from strategic tech conglomerates Tencent and Alibaba alongside parent firm XPeng Inc

    1 min