An inexpensive all-round model: text/photo to video, frame transition, extension and reference mode with audio. Handles prompts in any language.
PixVerse V6 is an affordable all-round model with the widest set of modes among budget options: ordinary text/photo generation, transition between two frames, extending an already-generated video, and a reference mode with named references.
💡 Useful tips:
Generate video from a text description — subject, action, location, camera, light. All 8 aspect ratios are available.
Upload 1-2 photos — the model will bring them to life. Describe only the motion, not the appearance (it's already visible in the photo).
Upload the first and last frame — the model creates a smooth, physically plausible transformation from one into the other.
Extend one of your own completed PixVerse videos — describe only the next action, without retelling the beginning.
Upload up to 7 references (characters or background) and assign each a role — subject or background. Refer to them in the prompt via @name: "@dog plays with @cat in @room". All 8 aspect ratios are available.
The official order: subject/identity → one action → location → one dominant camera movement → light/composition/materials → scene development → dialogue/audio (if enabled).
Words like "cinematic", "epic", "beautiful", "high quality" mean nothing to the model. Replace them with observable details: light direction, material behavior, physical signs of motion.
"Beautiful cinematic high quality video"
"Warm light from a window on the left, water splashing from under the wheels, the coat's fabric swaying with inertia"
Each reference has a role (subject/background) and an optional name. Use that name in the prompt EXACTLY as "@name" — without changing case or endings. Don't describe the reference's appearance — the model takes it from the photo.
Use only ONE dominant camera movement per shot (push-in, pull-out, pan, tracking, orbit, aerial, handheld, fixed). Combining several movements produces a jerky result.
Price = price per second × duration. The reference mode costs more than the ordinary ones (text/photo/transition/extend).
Study these prompts to see the structure behind effective descriptions and create better videos
POV shot from a bee's perspective, fisheye lens distortion, flying low across a suburban backyard toward a house with an open door. Wings blur with rapid motion, other bees swarm in the background. Camera flies forward continuously at insect height. Warm afternoon sunlight, shallow depth of field.
Giant armored kaiju monster smashes through a ruined city street, glowing yellow eyes, debris and concrete chunks flying past camera. Low-angle tracking shot from street level, skyscrapers crumbling on both sides. Overcast sky, dust and smoke haze, dramatic scale contrast.
Text, photo, frame transition, video extension, or reference — depending on the task.
Resolution (360p-1080p), duration (3-15 sec), audio, and aspect ratio (for text and reference modes).
PixVerse V6 handles prompts in any language well. Our AI can help improve your prompt using the official formula.
Click "Generate" — the result is usually ready in 1-5 minutes.
PixVerse V6 is an affordable all-round video generation model with five modes: text-to-video, image-to-video, frame transition, video extension, and reference-to-video with named references.
360p, 540p, 720p and 1080p. 360p is the cheapest option for drafts, 1080p for the final result.
Upload up to 7 reference images (characters or background) and refer to them in the prompt via @name, e.g. "@dog plays with @cat in @room". The model takes their appearance from the photo.
Upload the first and last frame — the model creates a smooth transformation from one frame into the other with physically plausible intermediate stages.
Yes, the "Extend" mode takes one of your own completed PixVerse videos and generates the next segment — just describe what should happen next.
You can write prompts in any language — PixVerse V6 handles them well. Translating to English is optional but usually makes results more stable.
Cost = price per second × duration (3-15 sec), rounded to a whole credit. The per-second price depends on resolution, audio, and mode (reference costs more than the ordinary modes).
Moderation is noticeably lighter than top-tier models, but it isn't absent — clearly prohibited content is still blocked.
5 modes, 3-15 seconds, from 360p to 1080p — from 3 credits/sec
We use cookies to operate the service, keep your session, and collect anonymous statistics. See our Privacy Policy.