Seedance 2.0 Quick Start Guide
Start from zero with Seedance 2.0 and complete your first multimodal video generation in 5 minutes.
Preparation
Before you begin, make sure you have registered a Seedance 2.0 workbench account. During the early launch period, high user volume may require queueing — off-peak hours are recommended.
Step 1: Choose Input Modalities
Seedance 2.0 supports any combination of four input modalities:
- Text instructions: Describe the visuals, camera movement, and mood you want
- Image references: Upload up to 9 reference images to control composition and style
- Video references: Upload up to 3 videos to reference camera work and action rhythm
- Audio references: Upload up to 3 audio clips to reference sound effects and music style
Step 2: Configure Output Parameters
Choose output specs suited to your use case:
| Parameter | Options |
|---|---|
| Duration | 4 / 8 / 12 / 15 seconds |
| Resolution | 480p / 720p / 1080p |
| Aspect ratio | 1:1, 21:9, 4:3, 3:4, 16:9, 9:16 |
Step 3: Write Director-Level Instructions
Good instructions should include:
- Subject description: Who/what is in the frame
- Action details: Specific movements
- Camera language: Push, pull, pan, tilt, track, follow, etc.
- Lighting and mood: Time, weather, color tone
Example: “A woman in a red dress walks slowly through rainy Tokyo streets. The camera slowly pushes in from a wide shot to a medium shot. Neon lights reflect on wet pavement. Soft jazz plays in the background.”
Step 4: Generate and Iterate
After clicking Generate, Seedance 2.0 outputs video and audio in sync. If results aren’t satisfactory:
- Refine text instruction precision
- Swap or add reference assets
- Use video editing for localized changes
- Use video extension to lengthen duration
FAQ
Q: How long does generation take?
A: Usually 1–3 minutes; longer during peak times.
Q: Does it support Chinese lip sync?
A: Yes. Seedance 2.0 natively supports lip sync in 8+ languages.
Happy creating!