
Alibaba has released the beta of its Wan3.0 video generation model, enabling creation of clips up to...
The AMW Read
Wan3.0 doubles video length and expands input modalities, advancing the foundation-model frontier for generative media, but it's an incremental update to an existing player.
Alibaba has released the beta of its Wan3.0 video generation model, enabling creation of clips up to 30 seconds per prompt—double the industry norm and its predecessor's limit of 15 seconds. Available via Alibaba Cloud's Model Studio and Qwen Cloud, the model accepts text, images, video, audio, and documents like webpages, PDFs, and PowerPoint files as references, and includes features for recommending optimal video length and extending narrative sequences.
This release signals a competitive escalation in generative video, where Chinese labs are pushing boundaries on duration and multi-modal input. Wan3.0's focus on reducing visual artifacts, rendering facial expressions, and supporting multilingual voice output positions it for film, short-form drama, social media, and corporate education—markets currently served by players like Kling and Runway. The ability to convert static documents into video content could also affect enterprise workflows where video production was previously cost-prohibitive.
For builders, the 30-second capability lowers barriers for narrative-driven content in marketing and training, enabling more coherent storytelling without post-production stitching. Investors should watch how Alibaba bundles Wan3.0 with its cloud ecosystem to drive enterprise adoption, potentially reshaping the competitive landscape for AI video tools.