Wan 3.0 Video

Alibaba's Wan 3.0 — one model for text-to-video, image-to-video, reference-to-video and video editing, with native audio and clips up to 30 seconds.

Coming Soon
Image To Video

Wan 3.0 Video is coming to Segmind

The playground and serverless API for this model aren't live yet. Here's what we know so far — launch details, pricing, and full API docs will appear on this page.

About

Wan 3.0 is Alibaba's all-in-one video generation model, and the biggest jump the Wan family has taken since 2.x. Where earlier releases split the work across separate text-to-video, image-to-video and reference-to-video models, Wan 3.0 is a single model that reads whatever you give it — a prompt, a first and last frame, reference images, reference clips, reference audio — and infers what you are asking for. One endpoint covers text-to-video, image-to-video, reference-to-video and video editing.

The headline change is length. Wan 3.0 generates up to 30 seconds in a single pass, double its predecessor, which is long enough for continuous shots and multi-beat camera moves instead of the two-to-five second fragments most models cap out at. It also generates a native audio track alongside the video, with multilingual voice output and voice consistency carried over from reference audio.

Alibaba took Wan 3.0 into public beta in August 2026. It is coming to Segmind as a single slug, wan3.0-video, with the same request shape across every mode.

What to expect

  • One model, four modes — text-to-video, image-to-video, reference-to-video and video editing, selected automatically from the media you attach
  • Up to 30 seconds of video in a single generation, or -1 to match the length of your source footage
  • Native audio, generated with the video rather than dubbed on afterwards, with multilingual voice output
  • 480P, 720P and 1080P output
  • Rich reference inputs — first frame, last frame, up to 10 reference images, and reference video or reference audio clips
  • Six aspect ratios plus an adaptive mode that follows your input
  • A single Segmind endpoint for every mode — no slug juggling between text, image and reference workflows

Specifications

ModelWan 3.0
ProviderAlibaba Cloud (Model Studio)
ModesText-to-video, image-to-video, reference-to-video, video editing
Resolution480P, 720P, 1080P (default 1080P)
Duration2–30 seconds; -1 matches the source length. With video input, input + output must stay within 30 seconds
Aspect ratioadaptive (default), 16:9, 4:3, 1:1, 3:4, 9:16
AudioNative audio track, on by default
Reference imagesUp to 10
Reference videoUp to 5 clips, 15 seconds total
Reference audioUp to 5 clips, 15 seconds total
First / last frame1 image each
StatusPublic beta at the provider

Upstream, Wan 3.0 also accepts documents and web pages — PDF, PPT, DOC, XLS, Markdown and more, up to 100 MB or 50 pages — as a video source. Document-to-video is not part of the initial Segmind release.

Availability

Wan 3.0 entered public beta on Alibaba Cloud Model Studio in August 2026, and the Segmind integration is built and in testing. The playground and the API go live on this page — check back here, and the Coming Soon badge will be replaced by a working playground, pricing and API docs the moment it ships.

Sources