ByteDance's flagship model: coherent scenes up to 30 seconds long, a huge reference budget and native stereo audio — all in a single generation, no stitching required.
Twice as long as Seedance 2.0. A full commercial or a short story in one piece — no stitching separate clips together, no character drifting between them.
Up to 30 images, 10 videos (30 seconds total) and 10 audio tracks in a single request. Assign each reference a role right in the prompt: @image1 for the character's face, @video1 for the camera move.
Sound effects, music, dialogue and voice-over are generated in the same pass as the picture, so they land on the action. Lip-sync works well in English and Russian alike.
The model handles long scenes better when you split them into time windows: 0-3s, 3-8s, 8-14s. Each window gets its own action and its own camera move — that is editing inside a single generation.
Text-to-video, first-frame animation, first-to-last-frame transition, last frame only, and multi-reference. The modes are mutually exclusive — pick one per generation.
Output format is chosen under "Advanced". MP4 plays everywhere; MOV is easier to pull into an editing timeline — but some browsers will make you download it instead of playing it inline.
Could not load pricing. Please refresh the page.
Describe the scene by the formula: subject → action → setting → visual style → camera work → sound. The model reads a prompt as a director's brief, not as prose.
For clips of 8 seconds and longer, split the scene into time windows (0-3s, 3-8s, 8-14s) with no gaps, and name a camera move in each one. That keeps the pacing tight and stops the scene from drifting.
Give every reference an explicit role: "@image1 defines the character's face and hair; do not take the background from it". If several photos show one object from different angles, say so directly.
Use one camera move per window. Combining a pan, a zoom and a tilt at once produces jittery footage. Changing the camera between windows is fine — that is the edit.
Phrase things positively: "smooth stable motion, natural proportions" instead of "no shake, no distortion". The model does not process negations and has no separate negative prompt.
Duration first: 2.5 generates up to 30 seconds versus 15 for 2.0. Second, the reference budget: 30 images, 10 videos and 10 audio files instead of 9 + 3 + 3. It also adds an output format choice (MP4 or MOV) and follows complex multi-shot prompts noticeably better. Version 2.5 costs more — for a short single shot, 2.0 is the cheaper pick.
Pricing is per second and depends on resolution. The cheapest option is 480p for 4 seconds, the most expensive is 1080p for 30 seconds. If you upload video references the per-second rate is lower, but it is multiplied by "reference seconds + result seconds". Exact numbers are in the table above and update automatically.
Yes. Dialogue and voice-over work well in English and Russian — you can write the lines directly. Chinese, Japanese, Korean, Spanish, Indonesian and Portuguese are supported as well. Audio is generated in the same pass as the picture, so it lands on the lips and on the action.
16:9, 9:16, 1:1, 4:3, 3:4, 21:9 and an "Auto" mode where the model picks the ratio to suit the scene. Auto is the default and usually gives the best result when you work with references.
No. Three scenarios are mutually exclusive: animating a first frame (optionally with a last frame), image multi-reference, and video multi-reference. Audio references can be added to the multi-reference modes. The form will tell you if a combination is not allowed.
MOV is a QuickTime container and is easier to pull into a professional editing timeline. For browser playback and social posting MP4 is better — it plays everywhere. When in doubt, leave MP4: it is the default.
It is a separate button in the form: you describe an idea, an AI director breaks it into shots and draws a single grid image containing all of them, then Seedance 2.5 animates that grid into one coherent scene with transitions between shots. Handy for multi-shot narrative and previsualisation.
Describe the idea — the AI turns it into a director's brief, and Seedance 2.5 shoots the scene with sound. From — credits.
We use cookies to operate the service, keep your session, and collect anonymous statistics. See our Privacy Policy.