Advanced Tips 2 min read Seedance Team
Multimodal Input Tips: Make AI Understand Your "Director Intent"
Master combination strategies for text, image, audio, and video modalities to convey creative intent precisely.
Multimodal Is Not Simple Stacking
Seedance 2.0 multimodal input is not “more assets equals better.” What matters is giving each modality a clear role. A strong multimodal prompt is like a concise storyboard.
Four-Modality Division Model
| Modality | Best Use | Recommended Count |
|---|---|---|
| Text | Narrative logic, detail, emotional tone | 1 concise instruction |
| Image | Visual style, character look, scene composition | 3–5 images |
| Video | Camera movement, action rhythm, transition style | 1–2 clips |
| Audio | Music style, sound character, speech rhythm | 1 clip |
Scenario 1: Brand Advertising
Goal: Generate a 15-second product showcase video
- Images: 3 high-res product shots (front, side, in-use scene)
- Video: Reference pacing from an existing brand ad
- Audio: Reference brand ad soundtrack
- Text: “Minimal premium style — product slowly lights up from darkness, 360° rotation, final hold on logo”
Scenario 2: Narrative Short
Goal: Generate a 12-second story-driven clip
- Images: 2 character sheets + 2 scene concept images
- Video: Reference cinematic camera work from a film clip
- Text: Full scene description with setup, development, and payoff
Scenario 3: Music MV Segment
Goal: Audio-visual sync for dance/performance
- Audio: Target song segment (model reads rhythm)
- Video: Reference dance moves or performance style
- Images: Stage/scene visual references
- Text: Supplement lighting and atmosphere
”Director Mindset” for Prompts
- Set the tone first: What emotion does this video carry?
- Define structure: How many shots? How long each?
- Then details: Specific content per shot
- Finally style: Visual and audio character
Debugging Tips
If output drifts from expectation, troubleshoot in this order:
- Check for contradictions between modalities
- Reduce reference count and add back one by one
- Split text instructions into more specific descriptions
- Use video editing to fine-tune instead of full regeneration