Getting Started 2 min read Seedance Team

Seedance 2.0 Quick Start Guide

Start from zero with Seedance 2.0 and complete your first multimodal video generation in 5 minutes.

Preparation

Before you begin, make sure you have registered a Seedance 2.0 workbench account. During the early launch period, high user volume may require queueing — off-peak hours are recommended.

Step 1: Choose Input Modalities

Seedance 2.0 supports any combination of four input modalities:

  • Text instructions: Describe the visuals, camera movement, and mood you want
  • Image references: Upload up to 9 reference images to control composition and style
  • Video references: Upload up to 3 videos to reference camera work and action rhythm
  • Audio references: Upload up to 3 audio clips to reference sound effects and music style

Step 2: Configure Output Parameters

Choose output specs suited to your use case:

ParameterOptions
Duration4 / 8 / 12 / 15 seconds
Resolution480p / 720p / 1080p
Aspect ratio1:1, 21:9, 4:3, 3:4, 16:9, 9:16

Step 3: Write Director-Level Instructions

Good instructions should include:

  1. Subject description: Who/what is in the frame
  2. Action details: Specific movements
  3. Camera language: Push, pull, pan, tilt, track, follow, etc.
  4. Lighting and mood: Time, weather, color tone

Example: “A woman in a red dress walks slowly through rainy Tokyo streets. The camera slowly pushes in from a wide shot to a medium shot. Neon lights reflect on wet pavement. Soft jazz plays in the background.”

Step 4: Generate and Iterate

After clicking Generate, Seedance 2.0 outputs video and audio in sync. If results aren’t satisfactory:

  • Refine text instruction precision
  • Swap or add reference assets
  • Use video editing for localized changes
  • Use video extension to lengthen duration

FAQ

Q: How long does generation take?
A: Usually 1–3 minutes; longer during peak times.

Q: Does it support Chinese lip sync?
A: Yes. Seedance 2.0 natively supports lip sync in 8+ languages.

Happy creating!