Wan lip sync prompting guide
Last checked 2026-10-11Wan 2.x / 3.0 (Alibaba) · Wan 2.5 Preview (t2v), Wan 2.6/2.7/3.0
How to prompt Wan in lip sync / dialogue mode. Every point links its source; community tips are marked. Your agent gets the same guide through the Atlas MCP (get_guide).
Prompt structure
image subject→audio track→style
- Entity + Scene + Motion; advanced adds aesthetic control (light, shot size, angle, lens, camera move) and stylization. source ↗
What the text should describe
- wan2.2-s2v: image_url + audio_url (wav/mp3, <15 MB, <20 s, clean speech without music); style speech|singing|performing; 480P/720P; portrait, half- or full-body, real or cartoon. Run wan2.2-s2v-detect on the image first. source ↗
What not to describe
- Don't rely on text-only prompts for word-accurate lip sync on Wan t2v/r2v. source ↗
Camera vocabulary
- Push-in for intimacy/tension, pull-out to reveal, tracking alongside, orbit (<45°) for importance, fixed for stillness. source ↗
Responds well to
- Voice = line + emotion + tone + speed + timbre + accent; SFX = source + action + ambient; orbit arcs under 45°. source ↗
Duration and limits
- Runway wan3: 2–30 s, 480p/720p/1080p, native audio; video refs ≤15 s total. wan2.7-videoedit input 2–10 s; wan2.2-s2v audio <20 s. source ↗
Multi-shot
- 'Shot 1 [0-3 s] ..., Shot 2 [3-6 s] ...'; on Wan 2.7 'Generate single shot.' forces one shot. source ↗
Avoid and failure modes
- Don't prompt: real people by name, rapid scene changes, exact text legibility, long choreographed action, lip sync to specific words. source ↗
Negative prompt
Supported via .
Next
All Wan modesWan model, prices & shotsFrom $0.05/s · pricesSeedance lip syncKling lip syncVeo lip syncRunway lip sync