Seedance Lip Sync (2026): How to Use It, Inputs & Limits

By SeedanceTips Team 19 min read

Seedance does not have a switch labelled “lip sync.” On the Seedance 2.0 series, speech and mouth movement are generated together with the picture, and you steer them in two ways: by writing the spoken line into the prompt inside curly braces, and by attaching a reference audio clip that supplies timbre, melody and dialogue content. That is the official framing on BytePlus ModelArk, and it is why “upload any face, upload any audio, get a talking head” does not describe what this model does.

The specs below were read from the official BytePlus ModelArk documentation on 2026-08-16. Where the docs are silent, this guide says so instead of filling the gap from third-party blogs.

What “lip sync” means on Seedance

The capability that produces synced speech is called omni reference. The ModelArk tutorial for the Seedance 2.0 series describes it verbatim as: “Input text, reference images, videos (with or without audio tracks), and audio to generate a new video. It can inherit core information including character image, visual style, and screen composition from reference images; subject, camera movement, action performance, and overall style from reference videos; and timbre, music melody, and dialogue content from reference audio.”

Three things follow from that sentence, and they are the whole mental model:

  1. Audio is a reference, not a soundtrack you bolt on. The model reads timbre, melody and dialogue content out of your clip and regenerates the performance. It is not muxing your file onto a rendered video.
  2. Dialogue can come from the prompt alone. You do not need an audio file to get a character speaking — the text prompt carries the line, and the model voices it.
  3. The face has to be one the model is allowed to use. Reference images and videos of real people are restricted (see the limits section).

The docs list four combined-reference modes for the Seedance 2.0 series, quoted from the capability table: “Image + audio”, “Image + video”, “Video + audio”, and “Image + video + audio”. Audio-driven work therefore always rides along with an image or video reference — there is no audio-only mode in that table.

Worth separating from all of this: Dreamina also ships a standalone tool called Lip sync, and it is a different product surface. Its page describes an AI-avatar workflow — verbatim, “Open Dreamina’s official website and navigate to the ‘Lip sync’ option on the homepage. Here, you can select the ‘AI avatar’ tab and click the ‘Import character image’” — after which you type the script or upload audio and pick a voice. That page never says the tool runs on Seedance and publishes no input limits, so none of the spec tables below apply to it.

Which channels support Seedance lip sync

Every cell below is either sourced or marked “Not stated.” Not stated means we could not find the answer on that channel’s own pages on 2026-08-16 — it is not a claim that the feature is missing.

ChannelSpoken dialogue written into the promptReference audio inputAudio-only reference (no image or video)
BytePlus ModelArk API — Seedance 2.0 seriesYesYes: up to 3 clips, 15 s totalNo — the 2.0 capability table lists audio only in combination with an image or a video
BytePlus ModelArk API — Dreamina Seedance 2.5YesYes: up to 10 clips, 30 s totalYes — the docs state “Seedance 2.5 newly supports generating videos with pure audio references, without requiring image or video assets”
Dreamina web app, Seedance 2.0YesYes — the tool page states up to 12 multimodal references: 9 images, 3 videos, 3 audio clipsNot stated
Dreamina “Lip sync” tool (AI avatar)Yes, as a typed scriptYes, via “Upload audio”Not applicable — a character image is required
Runway (Seedance 2.5)YesYes — Runway’s help centre describes a Reference mode taking up to 50 references: 30 images, 10 videos and 10 audio clipsNot stated
Other resellers (fal, Pollo and similar)Not statedNot statedNot stated

Two things follow. The audio envelope doubles between the 2.0 series and 2.5 — 15 seconds against 30 — so a job that will not validate on 2.0 may need the newer model rather than a re-cut. And the specs below are the model’s limits as published by BytePlus; a reseller can impose tighter ones, so check the vendor’s own page before sizing assets for a non-BytePlus channel.

Source for the audio and video limits in this section: BytePlus ModelArk, Video generation — reference and limitations, “Multimodal input” → “Audio requirements”. Read 2026-08-16.

Which Seedance model can actually do what

The three Seedance 2.0 series model IDs are dreamina-seedance-2-0-260128, dreamina-seedance-2-0-fast-260128 and dreamina-seedance-2-0-mini-260615, and the tutorial says they “support largely the same features, with the primary differences being the trade-off between generation quality and cost.” For dialogue work the difference that bites is resolution, not features: fast and mini top out below 1080p.

Before any of this works on BytePlus you also have to clear an account gate. The tutorial states the condition verbatim: “Before enabling the Dreamina Seedance 2.0 models, make sure you meet one of the following conditions: Recommended: BytePlus account balance > USD 30”, or a purchased AI Savings Plan at the USD 30 tier or above, or an active Seedance 2.0 resource pack with remaining quota.

Input requirements: image, audio and reference video

This is the part most third-party tutorials get wrong, usually by quoting Seedance 2.5 numbers for a 2.0 job or vice versa. Every figure in this table is from the ModelArk “Multimodal input” limitations section, read 2026-08-16.

Audio files

ItemSeedance 2.0 seriesDreamina Seedance 2.5
Input methodsAudio URL, Base64 string of audio, or asset IDSame
Formats.wav, .mp3.wav, .mp3
Length per clip2–15 seconds2–30 seconds
Number of clipsUp to 3Up to 10
Total duration of all clipsNo more than 15 secondsNo more than 30 seconds
File size15 MB per audio file15 MB per audio file
Request body ceiling64 MB64 MB

The docs add a practical warning next to the size limit: “Do not use Base64 encoding for large files.” Send a URL instead.

Reference video

ItemSeedance 2.0 seriesDreamina Seedance 2.5
Input methodsVideo URL or asset ID (no Base64)Same
Container formats.mp4, .mov.mp4, .mov
Video encodingH.264/AVC, H.265/HEVCH.264/AVC, H.265/HEVC
Audio encoding inside the fileAAC, MP3AAC, MP3
Length per video2–15 seconds2–30 seconds
Number of videosUp to 3Up to 10
Total durationNo more than 15 secondsNo more than 30 seconds
Resolution480p, 720p, 1080p, 4kSame
Aspect ratio (w/h)[0.4, 2.5]Same
Width and height300–6000 pxSame
Total pixel count407,696 to 8,295,044Same
File size200 MB per video200 MB per video
Frame rate24–60 fps24–60 fps

Reference image

ItemValue
Input methodsImage URL, Base64 string of image, or asset ID
Formats.jpeg, .png, .webp, .bmp, .tiff, .gif; Seedance 1.5 Pro and the 2.0 series also support .heic and .heif
Aspect ratio (w/h)(0.4, 2.5)
Width and height300–6000 px
File sizeUnder 30 MB per image; request body no more than 64 MB
Image count, first-frame image-to-video1 image
Image count, first-and-last-frame2 images
Image count, omni reference (2.0 series)1–9 images
Image count, omni reference (2.5)1–30 images

Output length and resolution

The duration parameter takes whole seconds: [4, 15] or -1 for the Seedance 2.0 series, [4, 30] or -1 for Dreamina Seedance 2.5. The docs explain the sentinel value: “A value of -1 enables intelligent duration selection, allowing the model to choose an appropriate video length within the supported range.”

Resolution matters for dialogue work because mouth detail is where low resolution shows first. resolution accepts 480p, 720p, 1080p and 4k, and ratio accepts 16:9, 4:3, 1:1, 3:4, 9:16, 21:9 and adaptive — but the docs carry two blunt caveats: “Dreamina Seedance 2.5, Dreamina Seedance 2.0 Fast, and Dreamina Seedance 2.0 Mini do not currently support 1080p or 4K output” and “Only Dreamina Seedance 2.0 supports 4K output.” A talking-head shot at 1080p therefore means the full dreamina-seedance-2-0-260128 model, not fast or mini.

How to turn lip sync on, step by step

On the BytePlus ModelArk API

There is no lip-sync flag to flip. You attach assets, each tagged with a role, and the model infers the task type. The Seedance 2.5 tutorial states the trigger condition verbatim: an omni reference-to-video task fires when “content contains at least one reference asset whose role is reference_image, reference_video, or reference_audio.”

  1. Clear the account gate (balance, savings plan or resource pack — see above) and get an API key.
  2. Choose the model. Use dreamina-seedance-2-0-260128 if you need 1080p or 4K; the fast and mini IDs are cheaper but capped lower.
  3. Build the content array. One text part carrying the prompt, then one part per asset. The official 2.0 sample uses {"type": "image_url", "image_url": {"url": "…"}, "role": "reference_image"} for stills, and the same shape with a video or audio type and the matching reference_video / reference_audio role for the other two.
  4. Put the spoken line in the prompt, using the bracket convention in the next section, and refer to each asset by index ([Audio 1], [Image 1]) inside the sentence that should use it.
  5. Set duration and resolution. duration is whole seconds, [4, 15] on the 2.0 series; -1 hands the choice to the model.
  6. Create the task and poll. Generation is asynchronous — the create call returns a task ID, not a video. Our Seedance API guide has the endpoint, the polling loop and the per-clip costs.
  7. Download promptly. The retention section of the 2.5 documentation states “Task records: Retained for 7 days”, and generated media URLs expire well before that, so re-host anything you intend to keep.

One asset-budgeting note: the prompt guide does not want you to max out the reference slots. Its recommended configuration is “1-2 character images (facial close-up / full body) + 1 scene image + 1 camera movement video + 1 audio clip”, warning that “Too many assets will make it difficult for the model to judge feature priorities.” For 2.5 it adds that subject audio and video references of 5-10 seconds generally work better than longer ones.

On Dreamina

Dreamina’s Seedance 2.0 tool page states the reference budget — 12 multimodal references, made up of 9 images, 3 videos and 3 audio clips, which lines up exactly with the ModelArk limits — but it does not publish a click-path for attaching an audio reference or a dialogue line. The step-by-step UI flow for Seedance audio references on Dreamina is not stated in official documentation as of 2026-08-16, so this guide does not invent one.

The separate Lip sync tool does have published steps, quoted verbatim from Dreamina’s page: “Go to the Lip sync option and enter the text you want your character to say. Besides, you can click the Upload audio option to upload the audio script. Choose the AI voice that suits your character. Finally, click the Generate button.” Again — that is the avatar tool, not the Seedance model.

How to write the prompt so the character speaks your line

The syntax is not the same on both model generations, which is why the advice you find online contradicts itself. The Seedance 2.0 series prompt guide defines an explicit four-symbol convention; the Seedance 2.5 prompt guide’s own worked examples use a different, plainer form. Take the one that matches the model you are calling.

For the 2.0 series, here is the “Special character standards” table, reproduced verbatim:

Information typeSymbolOfficial example
Music()(fast-paced rock music is playing in the background)
Sound effect<>< dog barking can be heard in the distance >
Dialogue{}{Hello, world}
Subtitles【】【Chapter One: Departure】

So the spoken line goes in curly braces, and the same table adds a rule for non-English, non-Chinese speech: “If the dialogue is in a less common language (neither Chinese nor English), the language must be marked, such as: says in Japanese {こんにちは}.”

Two more prompt rules come straight from the guide:

  • Do not code-switch inside a line. Verbatim: “The language of dialogue must be consistent, and mixing Chinese and English should be avoided (except for proper nouns).”
  • Name the voice when you attach reference audio. The guide’s worked fix for a mismatched voice is a prompt phrased like “Use the low, thick, warm, and finely grainy middle-aged male voice of @Audio 1 to say” — the reference asset is addressed by index inside the prompt text, and the voice is described in words as well as supplied as a file.

For Seedance 2.5, the official prompt guide’s examples drop the braces and write dialogue as a labelled line inside each shot — the pattern is Dialogue (elderly woman): "Fly safe, my child. Come back to me." sitting under a shot header such as Shot 3 (6-10s): Facial close-up. The 2.5 tutorial’s multilingual sample goes further and instructs the model directly, listing eight lines under the heading Lyrics (the lead vocalist sings the following "hello" in each language in order, with precise lip sync) — English, Chinese, Japanese, Korean, Portuguese, Thai, Spanish and Arabic — and repeats “precise lip sync” as a constraint at the end of the prompt. That example is BytePlus’s own, so “ask for precise lip sync in words” is a documented technique on 2.5, not folklore.

One more generation difference the 2.5 guide names outright: “Seedance 2.0 does not respond to timestamps and only responds to shot numbers, while Seedance 2.5 supports integer-second timestamps.” So on 2.0 you pace dialogue with Shot 1, Shot 2; on 2.5 you can write (6-10s) and expect it to land. The 2.0 omni-reference demo prompt still uses time ranges as prose (“2–4 seconds: …”) while addressing assets by index — [Video 1], [Image 1], [Audio 1] — so keep the indexing and treat the seconds as description rather than a timing contract.

Our Seedance prompt library has more prompt scaffolds, and the Seedance API guide covers the request structure these prompts sit inside.

Documented limits and what the docs do not say

Real human faces are restricted. The ModelArk limitations page states, verbatim: “Dreamina Seedance 2.5 and Dreamina Seedance 2.0 series models do not support direct uploads of reference images or videos containing real human faces.” This is the single biggest constraint on lip-sync work, because a talking head is usually a person. BytePlus documents three routes around it, quoted from its portrait-video page:

  • Trust model outputs as input assets. “The original face-containing outputs generated by certain models under this account can be used as input assets to call Seedance 2.5 and Seedance 2.0 series models again for secondary creation, without triggering input moderation blocking.” The trust window is 30 days from generation, it covers face-containing videos from Seedance 2.5 and the 2.0 series (effective from March 11, 2026), their last-frame images and Seedream 5.0 lite text-to-image outputs (both effective from April 16, 2026). The caveats are strict: same account only, same platform only, originals only — “they cannot be used after secondary editing or after the validity period has expired”, and “Compressing or forwarding files may invalidate trust verification.”
  • Use preset digital characters. A platform-provided library of “free, compliant, and diverse portrait assets”, for work that needs a realistic human face but not one particular person.
  • Use authorized real-person assets, which the page describes simply as support for “using authorized real-person portrait assets for video generation.”

For a talking-head series the workable pattern is therefore: generate the character once, keep working from your own recent outputs to stay inside the 30-day window, and store the originals unmodified.

Speech-specific failures the guide itself names. The Seedance 2.0 series prompt guide’s FAQ lists, among others, “Inaccurate Chinese pronunciation” — where the model “is prone to mispronouncing polyphonic characters, uncommon characters, and characters with similar shapes” — and “Inaccurate voice reference,” where “the audio voice in the final generated video differs significantly from the reference voice.” It also lists “Noise at the end of the video” and “Unexpected subtitles appear in the video.”

What the docs do not state. As of 2026-08-16 we could not find an official page giving any of the following, so this guide leaves them blank rather than borrowing a number from a third-party site:

  • A maximum number of spoken words per clip, or per line of dialogue. Official docs do not state this as of 2026-08-16. (Third-party prompt guides circulate figures like “20 words per 15 seconds”; none of them cite a BytePlus page, and we could not find one.)
  • A published list of languages the model can speak with synced mouth movement. The prompt guide only distinguishes Chinese, English and “less common” languages that must be labelled. Official docs do not state a supported-language list as of 2026-08-16.
  • A maximum number of simultaneous speaking characters in one shot. Official docs do not state this as of 2026-08-16.
  • A stated accuracy or quality figure for lip synchronisation. Official docs do not state this as of 2026-08-16, and this site does not publish effectiveness verdicts it has not tested.

Troubleshooting checklist

Work down this list before regenerating; most rejected or disappointing dialogue jobs fail on one of the first three lines.

  1. Clip too long or too many clips. On the 2.0 series each audio clip must be 2–15 seconds, no more than 3 clips, 15 seconds total. A single 20-second voiceover is out of spec — cut it, or move to Seedance 2.5’s 30-second envelope.
  2. Wrong container. Audio must be .wav or .mp3; video must be .mp4 or .mov with H.264/H.265 video and AAC/MP3 audio. A .m4a voice memo is not on the list.
  3. Base64 for a big file. The docs cap the whole request body at 64 MB and explicitly advise against Base64 for large files. Host the asset and pass a URL.
  4. Real face in the reference. See the restriction above; this is a hard model-level rule, not a moderation coin flip.
  5. Line in the wrong brackets. Dialogue is {}. Round brackets mean music and angle brackets mean sound effects, so a line written in parentheses may come back as scoring rather than speech.
  6. Voice drifts from your reference clip. The official fix is to describe the voice in words in the prompt as well as attaching the audio, and to keep the written line close in tone to the reference clip.
  7. Resolution ceiling. If you asked for 1080p or 4K and got neither, check which model you called: fast and mini do not do 1080p, and only the full Seedance 2.0 model does 4K.

FAQ

Is there a lip sync button in Seedance? No. On the Seedance 2.0 series and 2.5 the capability is called omni reference, and it activates when your request carries a reference asset or a spoken line in the prompt. The standalone button called Lip sync belongs to Dreamina’s AI avatar tool, which is a separate product surface.

What audio files does Seedance accept? .wav and .mp3, passed as a URL, a Base64 string or an asset ID. On the Seedance 2.0 series each clip must be 2-15 seconds, with at most 3 clips totalling 15 seconds. On Dreamina Seedance 2.5 each clip may be 2-30 seconds, with at most 10 clips totalling 30 seconds. Either way, a single file must stay under 15 MB and the whole request body under 64 MB. Data checked 2026-08-16.

Can I upload a photo of a real person and make them talk? Not directly. BytePlus states that the Seedance 2.5 and 2.0 series models do not support direct uploads of reference images or videos containing real human faces. It documents three alternatives: reusing face-containing output your own account generated in the last 30 days, using the preset digital character library, or using authorized real-person assets.

How do I write the spoken line in the prompt? On the Seedance 2.0 series, in curly braces — the prompt guide’s own example is {Hello, world}, and a non-English, non-Chinese line must be labelled with its language. On Seedance 2.5, the official examples instead write Dialogue (character): "line" under a shot heading. Match the convention to the model you are calling.

How many languages can Seedance speak with synced mouth movement? Official docs do not state a supported-language list as of 2026-08-16. The 2.0 prompt guide only distinguishes Chinese, English and “less common” languages that must be labelled; a BytePlus sample prompt for 2.5 demonstrates eight (English, Chinese, Japanese, Korean, Portuguese, Thai, Spanish, Arabic) but does not present that as the supported set.

Can Seedance generate a talking video from audio alone? On Seedance 2.5, yes: the docs say it “newly supports generating videos with pure audio references, without requiring image or video assets.” On the 2.0 series, no — its combined-reference table pairs audio with an image, a video, or both.

New to the model itself? Start with our Seedance 2.0 guide, then come back for the dialogue mechanics.

Sources

Every figure above was read on 2026-08-16 from:

SeedanceTips is an independent publication. It is not affiliated with ByteDance, BytePlus or CapCut, and this article contains no affiliate links.