Video Generation
Bring your generated images to life with video generation! Create videos right on the Generate page with reference media, frame control, and up to 4K resolution.
How Video Generation Works
Video generation takes a generated image and animates it based on your prompt, creating a short video clip with motion, effects, and life.
The Process
Start with an Image: Switch to Video on the Generate page, or tap the video icon on any generated image — from the Generate page gallery or your Collection page
Choose Content Mode: Advanced (Quality or Fast), Standard, or Adult
Set Your Frames or References (Advanced only): Pick a first frame, an optional end frame, or blend multiple reference photos and clips
Enter Video Prompt: Describe the motion and effects you want
Generate Video: AI creates a video based on your image and prompt
View Result: Watch your animated video in the gallery

Content Modes & Settings
Video generation supports three content modes:
Advanced: The primary video generation model, with reference media, frame control, and audio always on. Comes in two variants:
Quality: The highest fidelity results, with resolutions up to 4K
Fast: Quicker generations, with resolutions up to 720p
Standard: The legacy SFW model with optional audio
Adult: NSFW video generation (720p, 8 seconds, no audio)
Advanced (Quality)
480p, 720p, 1080p, 4K
4–15 seconds
Always on
Advanced (Fast)
480p, 720p
4–15 seconds
Always on
Standard
480p, 720p, 1080p
4–12 seconds
Optional
Adult
720p
8 seconds
Not available
Reference Media & Frame Control
Advanced mode gives you two ways to control your video, switchable via the Frames / References tabs:
Frames
Control exactly where your video starts and ends:
First frame: The image your video starts from. Set only a first frame and the model animates it based on your prompt.
End frame (optional): Add an end frame and the video will transition from your first frame to your last frame — great for reveals, outfit changes, and camera moves with a clear destination.
References
Blend multiple pieces of media to steer the generation:
Up to 5 reference photos (your source image counts as the first one)
Up to 3 reference clips, with a combined length of up to 15 seconds
The model uses your references to guide appearance, style, and motion rather than starting from a single fixed frame.
Picking Reference Media
The reference picker pulls from your entire gallery — every image and video you've generated. Filter by companion, tags, favorites, or most recent to find exactly what you need.
Video Prompting Guide
The Advanced video model is a massive leap in quality. It handles complex motion, realistic physics, multi-subject interactions, and synchronized audio generation all in a single pass. To get the most out of it, follow these prompting principles.
Core Prompt Structure
Since the image already provides the visual context, build your prompt around motion and change:
Action & motion — What moves and how (the most important element)
Camera movement — How the "camera" moves (one movement per shot)
Pacing — How fast or slow things happen
Audio cues — Dialogue, sounds, or atmosphere you want to hear
Constraints — What to avoid (no distortion, no jitter, etc.)
Key Principles
1. Focus on motion, not the scene
The image is your scene. Don't re-describe the setting, outfit, or appearance — the model already sees all of that. Instead, describe what changes.
✅ "She slowly turns her head and smiles, hair swaying gently in a breeze"
✅ "Leans forward slightly, eyes narrowing with a playful expression"
❌ "A woman standing on a beach in a red dress at sunset" — the image already shows this
2. Be specific about intensity and pacing
Use degree words to control how movements feel. The model is very responsive to pacing cues.
✅ "Slowly brushes hair behind her ear"
✅ "Gentle breeze lifts her hair, subtle movement"
✅ "Quickly glances over her shoulder"
❌ "Moves" — too vague, unpredictable result
3. One camera movement per shot
The model handles camera movement well, but only when you give it one clear instruction. Combining multiple movements causes jitter.
✅ "Slow dolly push-in"
✅ "Smooth pan to the right"
✅ "Camera holds still" — perfectly valid, keeps the focus on subject motion
❌ "Dolly in while panning left and tilting up" — too many movements at once
Use pacing words like "slow," "smooth," or "gentle" rather than technical parameters.
4. Separate camera movement from subject movement
This is the most common prompting mistake. Be explicit about what moves and what stays still.
✅ "She slowly turns her head to the right. Camera holds fixed framing."
✅ "Camera slowly pushes in. She holds her pose, only her hair moves in the wind."
❌ "Everything moving at once" — confuses the model
5. Use sequential actions for complex scenes
List actions in the order you want them to happen. The model follows temporal sequences well:
"Looks down at her phone with a surprised expression, then looks up at the camera with excitement, gasps softly"
"Takes a sip of coffee, sets the cup down, then glances out the window with a thoughtful expression"
6. Add quality constraints at the end
Append a short constraint line to any prompt for consistently better output:
"No distortion. No jitter. Face stable, no deformation."
This works across all content types and noticeably reduces common artifacts.
Example Prompts
Simple & effective:
She brushes hair behind her ear with a soft smile, gentle breeze, smooth and natural motion. No distortion.
With camera movement:
Slow dolly push-in. She looks up toward the camera with a warm, inviting expression, lips parting slightly. Hair shifts gently. No jitter, face stable.
Complex sequential action:
She glances down at her phone and her eyes widen with surprise, then she looks up at camera with excitement and laughs softly. Camera holds still. No distortion, no deformation.


Prompting With Audio
Audio is generated alongside the video in Advanced mode — not as a separate step. The model reads the source image and your prompt to produce matching sound automatically.
Three audio layers are generated simultaneously:
Lip-synced speech — Include short dialogue in your prompt with emotional context for best results (e.g., "She softly whispers 'Hey, I missed you'"). Keep dialogue to 5–10 words per line for the cleanest lip sync.

Sound effects — Tied to the actions you describe. Footsteps, cloth movement, glass clinking, a sigh — the model generates these based on what's happening in the scene.


Ambient sounds & music — Background atmosphere inferred from the source image and prompt (city noise, ocean waves, quiet room tone, soft background music).
Things to Avoid
Re-describing the image
The model already sees the image — redundant prompting wastes space
Focus on motion and change
Multiple camera movements in one shot
Causes jitter and incoherent motion
One clear camera movement per shot
Contradicting the source image
Creates visual confusion (e.g., describing a beach when the image is indoors)
Stay consistent with what's in the image
Overloading a short clip
4–15 seconds can't contain an entire story
One clear moment per generation
Vague motion ("she moves")
Unpredictable results
Be specific about what moves and how
Long dialogue lines
Lip sync degrades with long speech
Keep dialogue to 5–10 words

Requirements & Limitations
Subscription Requirements
FREE Tier: ❌ Not available
PREMIUM Tier: ✅ Available
ULTIMATE Tier: ✅ Available
Understanding Video Results
What to Expect
The Advanced model handles complex motion, realistic physics, and multi-subject interactions far better than previous models. That said, AI video generation still has some limitations:
Hands & fingers — Occasional extra or fused digits can appear in complex hand movements
Fine patterns — Tight weaves, tiny checks, and detailed textures can shimmer or flicker. Simpler patterns work better.
Variation — Results vary between generations, even with the same prompt
Iteration — You may need to regenerate or tweak your prompt for the best result
Frequently Asked Questions
Q: How much do videos cost?
Advanced: Cost scales with resolution and duration — the exact price is shown on the Generate button before you generate. Audio is always on and included.
Standard: Cost scales with resolution, duration, and whether audio is enabled.
Adult: 600 moments (720p, 8 seconds).
Adding reference photos or clips costs nothing extra.
Q: Do I need a subscription?
Yes, video generation requires a PREMIUM or ULTIMATE subscription.
Q: Can I use any image?
Yes, any generated image can be used as source for video.
Q: What resolutions can I generate?
Advanced (Quality): 480p, 720p, 1080p, or 4K.
Advanced (Fast): 480p or 720p.
Standard: 480p, 720p, or 1080p.
Adult: 720p.
Q: How do I generate in 4K?
Choose Advanced mode and the Quality variant — 4K is exclusive to Advanced (Quality).
Q: What's the difference between Quality and Fast?
Quality produces the highest fidelity results and supports up to 4K. Fast generates quicker and tops out at 720p. Both support the full 4–15 second range, reference media, frame control, and audio.
Q: Can I use reference photos and clips?
Yes! In Advanced mode, switch to the References tab to blend up to 5 reference photos and 3 reference clips (15 seconds of clip time combined) from your gallery. It costs nothing extra.
Q: Can I control the first and last frame?
Yes! In Advanced mode, use the Frames tab to set a first frame, and optionally an end frame — the video will transition from one to the other.
Q: Can I add audio?
In Advanced mode, audio is always on. In Standard mode, audio is optional.
Audio supports speech lip sync, ambient sounds, action sounds, and background music.
Q: Can I download videos?
Yes, you can download generated videos from the videos tab.
Q: Can I generate multiple videos from same image?
Yes! Generate as many videos as you want from the same image with different prompts.
Q: What motion can I create?
The Advanced model excels at complex, realistic motion — walking, turning, hair blowing in wind, clothing movement, facial micro-expressions, multi-subject interactions, and much more. It handles physics naturally, so don't be afraid to get creative.
Q: How do I get the best results?
Start your prompt with the shot type, always describe the lighting, use one camera movement per shot, and add "No distortion. No jitter." at the end. See the Video Prompting Guide for detailed tips.
Want to learn more? Check out Image Generation or explore Custom Characters!
Last updated
Was this helpful?

