Up to 30 seconds in a single take, native sound and dialogue, up to 20 reference files at once — and the only mode on the site that turns a document or a link into video.
The AI assistant turns your idea into a full prompt with lighting, camera and sound, then offers to translate it into English — the language the model understands best.
The scene is built from scratch out of your description. On longer clips, split the prompt into shots with timecodes — the model holds the pacing far better.
The first frame sets the look and the location, the last one sets the ending. Describe only the motion and the camera work: re-describing the picture fights against it.
Up to 10 images, 5 videos and 5 audio files at once. Editing and extending an existing video happen here too, through the wording of the prompt.
PDF, DOCX, PPTX, XLSX, TXT, MD or a public web page. The model reads the material itself while you write the brief for filming it.
Standard rendering speed
Faster and pricier — same capabilities
Important: both the input videos and the result are billed. Fifteen seconds of reference video plus a fifteen-second clip is charged as thirty. Trimming the reference right in the form lowers the cost.
Dialogue, effects and music arrive in the same pass as the picture, and speech is lip-synced. If you do not need audio, switch it off and the model will not spend attention on it.
WAN 3.0 is the next generation of Alibaba's Wan-AI video model. It generates up to 30 seconds in a single take with native audio. Four modes: text to video, frames (first and last), references (up to 10 images, 5 videos and 5 audio files), and a one-of-a-kind document-or-link to video mode.
Clips are twice as long — up to 30 seconds instead of 15. There is room for far more references: 10 images, 5 videos and 5 audio files, against five items in total before. It adds 480P, vertical aspect ratios, and generation from a document or a link. WAN 3.0 has no negative prompt: anything unwanted is described in words inside the prompt itself.
Attach a PDF, DOCX, PPTX, XLSX, TXT or MD file — or paste a link to a public web page. The model reads the material itself and films a clip from it. Use the prompt for the brief: which part to put on screen, for which audience, at what pace. No other model on the site offers this.
The tiers have identical capabilities — only rendering speed and price differ. Prime is noticeably faster and about one and a half times more expensive. If you are not in a hurry, the standard tier is enough.
Per second, and both the input videos and the result are billed: cost = rate × (seconds of reference video + clip length). The rate depends on the tier and the resolution.
The model processes an uploaded video the same way it generates a new one, so its seconds are billed alongside the result. Fifteen seconds of reference plus a fifteen-second clip is charged as thirty. Trimming the reference right in the form genuinely lowers the cost.
The model simply has no such field. Anything unwanted is described in positive prose inside the main prompt: instead of "no watermark", write "a clean frame with no overlaid text or logos". Short functional negations such as "no background music" and "no subtitles" are understood literally and can go straight into the text. Our AI assistant adds them for you.
In the prompt, files are referred to as Image 1, Video 2, Audio 1 — capitalised, with a space, in English. Numbering follows upload order and is counted separately for each type. Give every file a single stated role — "Image 1 is the only appearance reference for the hero" — otherwise the model will mix up what comes from where.
There is no separate button — it happens in the References mode through the wording of the prompt. To edit: "Transform Video 1 into ..., everything else stays the same". To extend: "Extend Video 1 forward: ..." followed by a description of the continuation.
No. Write in your own language — the AI assistant will build the prompt and then offer a separate step to translate it into English, which the model understands best. Cinematography terms are left in English automatically.
Auto, 16:9, 9:16, 4:3, 3:4 and 1:1. In Auto the model picks the ratio from your uploaded files — handy when the references are already vertical or square.
Usually three to fifteen minutes: the longer the clip and the higher the resolution, the longer it takes. The Prime tier renders noticeably faster. You do not have to stay on the page — the finished video shows up in the list on its own.
We use cookies to operate the service, keep your session, and collect anonymous statistics. See our Privacy Policy.