How to Use Seedance 2.0: Step-by-Step Tutorial for Beginners (2026)
2026/04/07

How to Use Seedance 2.0: Step-by-Step Tutorial for Beginners (2026)

I tested 300+ Seedance 2.0 prompts so you don't have to. Complete tutorial with 15 copy-paste prompt templates, the @ reference system explained, and cinematic techniques — from first generation to multi-shot storytelling.

I spent my first week with Seedance 2.0 burning through credits on morphed faces, stiff walks, and scenes that looked like a cheap slideshow. The model was clearly powerful — but my prompts were the problem. After 300+ generations, I figured out what actually works: a three-layer prompt structure, the @ reference system most tutorials skip, and a handful of templates that produce cinematic results on the first try. This is the tutorial I wish I had on day one.

What Is Seedance 2.0 — 30-Second Overview

Seedance 2.0 is ByteDance's latest AI video generator, released in February 2026. If you have used Seedance 1.5, here are the three upgrades that matter:

  • Multimodal input: Feed it text + up to 9 images + 3 videos + 3 audio files simultaneously
  • Native audio: Sound effects, music, and multilingual lip-sync dialogue built into the generation pipeline
  • 15-second duration: Up from the previous cap, with multi-shot storytelling in a single generation
SpecDetails
Resolution480p, 720p, 1080p (Standard); 480p, 720p (Fast)
Duration4–15 seconds
Aspect Ratios16:9, 9:16, 1:1, 4:3, 3:4, 21:9
Input TypesText, images, video clips, audio tracks
AudioNative sync — sound effects, music, dialogue

For a deep dive into features and how Seedance 2.0 compares to Kling and Sora 2, read the Seedance 2.0 Complete Guide.

How to Access Seedance 2.0

You do not need a Chinese phone number or VPN. The fastest path is using SoraVideo.art's Seedance 2.0 tool directly in your browser — text-to-video, image-to-video, and reference-to-video all in one place.

Start generating in 60 seconds

Go to SoraVideo.art → Seedance 2.0 and start with text-to-video. No separate ByteDance account required. New users can grab a $4.99 starter pack — enough credits to test every template in this guide. If it fits your workflow, upgrade to a monthly plan later. For other access options including Dreamina and BytePlus API, see the full access guide.

Your First Video in 5 Minutes

Most tutorials only show text-to-video. That is like learning to drive only in first gear. Seedance 2.0 has three generation modes — pick the one that matches what you have:

  • Text-to-Video → You have an idea in your head but no assets
  • Image-to-Video → You have a photo and want it to come alive
  • Reference-to-Video → You have images + video clips + audio and want to mix them all

Here is the step-by-step for your first generation:

Start with Text-to-Video for your first attempt. Set these beginner-friendly parameters:

ParameterRecommendedWhy
Duration5 secondsFast iteration, lower credit cost
Resolution720pBest quality-to-speed balance
Aspect Ratio16:9Universal format
AudioOnSeedance's unique strength — don't waste it

Copy this starter prompt and paste it directly:

A golden retriever runs through a sunlit meadow toward the camera.
Slow motion. Cinematic golden hour lighting. Shallow depth of field.
Camera tracks the dog at eye level, slight handheld movement.
Warm color grading, film grain texture.

Why this works:

  • Subject is specific: "golden retriever" not "a dog"
  • Action is clear: "runs through a meadow toward the camera"
  • Camera has direction: "tracks at eye level, handheld movement"
  • Style is defined: "golden hour, shallow DOF, film grain"

Click Generate and wait 30–60 seconds. When the result appears, ask yourself three questions:

  1. Is the subject right? If the dog looks wrong, be more specific about breed, color, and size.
  2. Is the motion right? If the movement feels stiff, add words like "natural gait" or "energetic sprint."
  3. Is the camera right? If the framing is off, specify the distance — "medium shot" or "close-up."

Change one variable at a time. Do not rewrite the entire prompt — that resets the model's interpretation and wastes credits.

Your first video will not be perfect. That is fine. Iteration is the skill, not luck.

The Three-Layer Prompt Structure

After testing hundreds of prompts, I keep coming back to the same structure. Every effective Seedance 2.0 prompt has three layers:

LayerWhat to writeExample
SceneLocation, time, weather, atmosphererain-soaked neon alley at midnight, steam rising from grates
SubjectCharacter, action, expression, wardrobewoman in black trench coat, walking fast, tense expression
Camera & StyleLens, movement, color grade, film stocktracking shot from behind, film noir palette, 35mm grain

Rules I learned the hard way

The model reads left to right. What you write first carries the most weight. If your subject keeps getting overshadowed by the background, move the subject description to the very beginning of your prompt.

150 words is the sweet spot. Under 50 words and the model fills in the blanks randomly. Over 300 words and it starts ignoring the second half. I tested this extensively — 120 to 200 words consistently gives the best results.

Use positive descriptions, never negatives. Write "sharp focus, clean details" instead of "no blur, no artifacts." The model interprets positive instructions far more reliably.

Mix English and Chinese when it helps. Action descriptions work better in English. Mood and atmosphere keywords sometimes produce stronger results in Chinese. My personal approach: English for the main prompt, Chinese for style keywords when targeting a specific aesthetic.

The @ Reference System — Most Tutorials Skip This

This is the feature that separates Seedance 2.0 from every other AI video generator. Most tutorials mention it in one sentence. That is a mistake — it is the single most powerful tool for consistent, controllable results.

What you can upload

TypeLimitWhat it controls
Images0–9Character appearance, environment, style
Videos0–3 (total ≤ 15 seconds)Camera movement, motion style, pacing
Audio0–3 (total ≤ 15 seconds)Music rhythm, dialogue lip-sync, ambient sound

How the priority system works

The model weighs references in this order:

  1. @Audio = Rhythm anchor — drives lip-sync timing and beat-matched editing
  2. @Video = Motion anchor — copies camera trajectories and action choreography
  3. @Image = Visual anchor — locks character face, wardrobe, and environment style

Complete example prompt

Character from @Image1 walks through the landscape in @Image2.
Camera follows from @Video1 perspective, smooth tracking movement.
Background music synced to @Audio1, beat-matched cuts.
Cinematic warm lighting, shallow depth of field, 35mm film texture.

Practical tips

  • Start with 2–3 files, not 12. More references means more conflicting signals. Add complexity gradually.
  • For character images, use a mid-body portrait with a clean background. Busy backgrounds confuse the visual anchor.
  • Always specify what each reference controls in your prompt. Do not just upload files and hope — tell the model explicitly: "face from @Image1" or "camera movement from @Video1."

15 Prompt Templates You Can Copy Right Now

Every template below is ready to paste. I include the prompt, why it works, and the recommended parameters.

Text-to-Video Prompts

1. Product Commercial

A sleek espresso machine sits on a marble kitchen counter.
Steam rises from a freshly pulled shot. Morning sunlight streams
through a window, casting warm highlights on brushed steel.
The camera performs a slow 180-degree orbit around the machine.
Shallow depth of field, product photography lighting,
warm neutral color grading, 4K commercial quality.

Why it works: Specific product + natural environment + controlled camera orbit + commercial lighting language. Parameters: 16:9 · 10 seconds · 720p · Audio on


2. Anime Action Scene

Two samurai face each other in a bamboo forest at dawn.
Wind ripples their clothing. Leaves fall slowly between them.
In a sudden burst, both draw swords simultaneously.
Steel clashes with a brilliant spark. The camera whip-pans
between close-ups of their determined eyes and the
impact point. Anime cel-shading style, dramatic speed lines,
saturated color palette, dynamic composition.

Why it works: Clear setup → action beat → camera instruction + anime-specific style keywords. Parameters: 16:9 · 8 seconds · 720p · Audio on


3. Fight Scene Choreography

A martial artist in a white gi executes a spinning roundhouse kick
in an industrial warehouse. Dust particles explode from the impact
point. The camera captures the kick in slow motion from a low angle,
then snaps to real-time as the fighter lands and transitions into
a defensive stance. Dramatic side lighting with deep shadows.
Cinematic action film style, sharp focus on the fighter,
motion blur on the kick arc.

Why it works: Specific martial arts move + environmental reaction (dust) + speed change (slow-mo → real-time) + strong lighting direction. Parameters: 16:9 · 10 seconds · 720p · Audio on


4. Short Film Opening

A lone figure in a worn overcoat walks through an abandoned
train station at twilight. Broken glass crunches under each step.
Pigeons scatter from the rafters. Volumetric light shafts cut
through dusty air from shattered skylights. The camera begins
with a wide establishing shot, then slowly dollies forward,
following the figure deeper into the station.
Desaturated teal-and-amber color grade, anamorphic lens flares,
melancholic atmosphere, film grain.

Why it works: Sensory details (glass crunching, pigeons, dust) + classic dolly-forward opening + specific color grade + emotional tone. Parameters: 21:9 · 8 seconds · 720p · Audio on


5. Social Media Hook (Vertical)

Extreme close-up of latte art being poured into a ceramic mug.
Milk swirls into a perfect rosetta pattern. Steam rises gently.
The camera is locked overhead, shooting straight down.
Warm morning light from the left. Cozy café aesthetic,
soft bokeh background, ASMR-quality detail.
Slow motion, satisfying visual rhythm.

Why it works: Overhead lock = stable framing for vertical video + ASMR/satisfying content hook + specific latte art pattern. Parameters: 9:16 · 5 seconds · 720p · Audio on

Image-to-Video Prompts

6. Product Photo → Promo Video

Upload your product photo as the reference image, then use this prompt:

The product in @Image1 sits on a clean surface.
The camera slowly orbits 180 degrees around it.
Soft studio lighting highlights surface textures and materials.
A subtle reflection appears on the surface below.
Product photography style, premium aesthetic,
smooth continuous camera movement.

Why it works: @Image1 locks the product appearance + orbit camera shows all angles + studio lighting language triggers professional rendering. Parameters: 16:9 · 8 seconds · 720p · Audio off


7. Portrait → Character Animation

Upload a portrait photo, then:

The person in @Image1 turns their head slowly from left to right,
as if noticing something in the distance. A gentle breeze moves
their hair. Natural expression transitions from neutral to
a slight smile. The camera holds a medium close-up, steady.
Soft natural lighting, cinematic skin tones,
shallow depth of field on the background.

Why it works: Simple, controlled motion (head turn) + natural environmental interaction (breeze) + held camera avoids distortion. Parameters: 16:9 · 5 seconds · 720p · Audio off


8. Landscape → Cinematic Pan

Upload a landscape photo, then:

The landscape in @Image1 comes alive with subtle motion.
Clouds drift slowly across the sky. Grass sways in a gentle wind.
Light shifts as the sun moves behind a cloud.
The camera performs a slow horizontal pan from left to right,
revealing the full scene. Epic cinematic scale,
golden hour warmth, wide-angle lens perspective.

Why it works: Asks for environmental motion (clouds, grass, light) rather than character motion, which landscapes handle well. Parameters: 21:9 · 10 seconds · 720p · Audio on


9. First-Frame + Last-Frame Control

Upload two images — the model automatically uses the first as the opening frame and the second as the closing frame:

Smooth cinematic transition from @Image1 to @Image2.
The camera slowly pushes forward as the scene transforms.
Lighting transitions naturally between the two environments.
Maintain visual continuity and fluid motion throughout.
Dreamlike quality, gentle pace, soft color blending.

Why it works: Two-image upload triggers Seedance 2.0's first-last frame mode automatically. The prompt guides the transition style. Parameters: 16:9 · 6 seconds · 720p · Audio on


10. Consistent Character Across Shots

Upload your character reference image, then use this prompt for every generation in your project:

Character from @Image1 walks down a busy city street at night.
Maintain facial features and wardrobe fully consistent with @Image1:
short black hair, blue denim jacket, white sneakers.
Natural walking pace, confident posture.
The camera tracks from a side angle at waist height.
Urban night atmosphere, neon reflections on wet pavement.

Why it works: Double anchor — @Image1 visual reference + explicit text description of key traits. This redundancy is essential for consistency across multiple generations. Parameters: 16:9 · 8 seconds · 720p · Audio on

Reference-to-Video Prompts (Seedance 2.0 Exclusive)

These templates use the full multimodal capability that only Seedance 2.0 offers.

11. Multimodal Mix — Image + Video + Audio

Upload: character photo + reference video clip + background music track.

Use the first-person perspective framing of @Video1 throughout.
Use @Audio1 as background music throughout, beat-synced editing.
Character from @Image1 walks through a neon-lit street market.
Camera follows the character from behind, matching the movement
style in @Video1. The character pauses to examine a food stall,
turns to the camera, and smiles. Cinematic night photography,
rich saturated colors, shallow depth of field.

Why it works: Each @ reference has a clearly defined role — @Video1 for camera, @Audio1 for rhythm, @Image1 for character. Parameters: 16:9 · 10 seconds · 720p · Audio on


12. Video Editing — Replace Elements

Upload: original video + replacement product image.

Replace the object being held in @Video1 with the product
shown in @Image1. Keep the original camera movement,
lighting, and hand gestures unchanged.
Maintain natural interaction between the hand and the new product.
Seamless integration, matching color temperature and shadows.

Why it works: Tells the model exactly what to preserve (camera, lighting, gestures) and what to replace (the object). Parameters: Match original video aspect ratio · Match original duration · 720p · Audio on


13. Video Extension — Multi-Segment Continuation

Upload: 2–3 video segments you want connected.

Continue the narrative from @Video1 into @Video2.
Maintain consistent character appearance, lighting direction,
and color grade across the transition. The camera movement
flows naturally between segments without jump cuts.
Smooth temporal blending at transition points.
Cinematic continuity, professional editing feel.

Why it works: Explicit continuity instructions prevent the jarring style shifts that happen when stitching clips. Parameters: Match original aspect ratio · 8 seconds · 720p · Audio on


14. Music Video — Audio-Driven Generation

Upload: character photo + music track.

Character from @Image1 performs to the rhythm of @Audio1.
Movement intensity follows the music dynamics: subtle sway during
quiet sections, energetic motion during the chorus.
Camera cuts sync to beat drops. Lighting pulses with the rhythm.
Music video aesthetic, high contrast, dramatic color grading,
concert-style spotlights.

Why it works: Links specific visual behaviors (movement, cuts, lighting) to specific audio events (beats, chorus, quiet sections). Parameters: 9:16 · 15 seconds · 720p · Audio on


15. Talking Head — Lip-Sync Dialogue

Upload: character photo + audio dialogue recording.

Character from @Image1 speaks the dialogue from @Audio1.
Precise lip-sync matching the audio. Natural head micro-movements:
slight nods when emphasizing points, occasional blink,
subtle eyebrow raises. Professional presenter framing,
medium close-up, clean background.
Three-point studio lighting, warm skin tones, sharp focus on face.

Why it works: Specific micro-movement instructions (nods, blinks, eyebrow raises) prevent the "frozen face" problem in talking head videos. Parameters: 16:9 · Match audio duration · 720p · Audio on

Camera Language Cheat Sheet

You do not need to be a filmmaker. Just pick one camera term from this table and add it to your prompt. Instant upgrade.

TermWhat it doesWhen to use it
Dolly inCamera pushes forward toward subjectBuilding emotional intensity
Dolly outCamera pulls back from subjectRevealing context or environment
Tracking shotCamera follows subject laterallyAction sequences, walking scenes
Crane downCamera descends verticallyGrand establishing shots
HandheldSlight natural shakeDocumentary feel, urgency
360° orbitCamera circles around subjectCharacter reveals, product shots
Whip panUltra-fast horizontal camera snapTransitions between subjects
Rack focusFocus shifts from foreground to backgroundDrawing attention between elements
Dutch angleCamera tilted on its axisTension, unease, stylized shots
POVCamera represents character's eyesImmersive first-person perspective

Keeping Characters Consistent

Character consistency is the hardest part of AI video. Here is the system that works:

  1. Always use the same @Image1 across every prompt in your project. This is your visual anchor.
  2. Repeat key physical traits in text even when you have a reference image. Write "short silver hair, scar on right cheek, black leather jacket" every time. The model needs both visual and text anchors.
  3. Keep character design simple. Fewer accessories and simpler outfits produce more stable results across generations.
  4. Use a clean mid-body portrait as your reference image. Busy backgrounds or extreme angles confuse the visual anchor system.

Common Prompt Mistakes and Fixes

MistakeBad exampleFix
Too vague"a nice video of nature"Add specifics: "red fox crossing a frozen river at dawn, wide shot"
Contradictory instructions"fast-paced slow motion"Choose one: "slow motion" or "fast-paced editing"
Prompt too long300+ wordsCut to 120–200 words — front-load the important parts
No camera directionCompletely absentAdd at least one: "tracking shot" or "static medium shot"
Negative framing"no blur, no artifacts"Rewrite as: "sharp focus, clean output, crisp details"
Overloading references12 files uploaded at onceStart with 2–3 files. Add more only if needed

FAQ

What is the best prompt length for Seedance 2.0? 150 words is the sweet spot I keep coming back to. Under 50 and the model guesses too much. Over 300 and it starts ignoring the second half. Front-load the most important elements — subject and action first, style and camera second.

Can I use Chinese prompts in Seedance 2.0? Yes, and sometimes they produce stronger results for mood and atmosphere. I mix both — English for subject and action descriptions, Chinese for style keywords when targeting a specific aesthetic. Both languages work natively.

How do I keep characters consistent across multiple videos? Always use the same @Image1 reference. Then repeat key traits in text every time: "short silver hair, mole under left eye, blue jacket." The model needs both visual and text anchors. Simple character designs are more stable than complex ones.

What is the difference between Seedance 2.0 and Sora 2 prompting? Different tools respond to different prompt styles. Sora 2 reads atmosphere and emotion words more effectively. Seedance 2.0 responds better to specific action descriptions and @ reference instructions. I wrote a separate Sora 2 prompts guide if you want to compare approaches.

How long does generation take? A 5-second clip at 720p typically takes 30–60 seconds. A 15-second clip with audio enabled can take 90–120 seconds. Lower resolution (480p) generates faster if you are iterating on prompt language.

Can I use Seedance 2.0 videos commercially? Licensing varies by platform. On SoraVideo.art, generated content is yours to use commercially. Check each platform's specific terms before publishing commercial content.

Does Seedance 2.0 support negative prompts? Not as a separate field like image generators. Instead, use positive framing: "sharp focus" instead of "no blur," "stable composition" instead of "no shaking." The model follows affirmative instructions more reliably.

The Bottom Line

Seedance 2.0 is the most capable AI video generator I have used in 2026 — but only if you know how to talk to it. The three-layer prompt structure, the @ reference system, and the 15 templates above are everything I use daily.

Start with one template. Generate. Review. Change one thing. Generate again. You will be producing cinematic clips within your first session.

New to Seedance 2.0? Read the access guide to get set up, then come back here for the prompts. Want to understand costs before scaling up? Check the Seedance 2.0 pricing breakdown. Curious about the full feature set? The Seedance 2.0 complete guide covers everything under the hood.

Start creating with Seedance 2.0

Access Seedance 2.0, Sora 2 Storyboard, Kling Motion Control, and more on SoraVideo.art — all your AI video tools in one place. See plans.

Newsletter

Join the community

Subscribe to our newsletter for the latest news and updates