Models / Grok Imagine / Lip sync
Grok Imagine lip sync prompting guide
Last checked 2026-10-11Grok Imagine Video 1.5 (xAI) · grok-imagine-video-1.5, 1.5 Lite, grok-imagine-video (edit/extend)
How to prompt Grok Imagine in lip sync / dialogue mode. Every point links its source; community tips are marked. Your agent gets the same guide through the Atlas MCP (get_guide).
Prompt structure
speaker ref→line→voice ref
- No dedicated prompting guide published; describe subject, camera, pacing and sound in plain language. source ↗
What the text should describe
- 'The person from <IMAGE_0> ... speaking with the voice from <AUDIO_0>.' Audio is on by default. source ↗
Responds well to
- t2v first generates a frame from your prompt then animates it, so the opening frame description matters. source ↗
Duration and limits
- Duration 1–15 s (default 8). 1.5 native 1080p for t2v/i2v; r2v up to 720p. i2v follows the image's aspect ratio. Editing keeps source length (max 8.7 s), 720p. source ↗
Negative prompt
Supported via .
Real prompts, credited
The person from <IMAGE_0> presents the product from <IMAGE_1> on the set from <IMAGE_2>, speaking with the voice from <AUDIO_0>. A second speaker with the voice from <AUDIO_1> replies.
Next
All Grok Imagine modesGrok Imagine model, prices & shotsFrom $0.02/s · pricesSeedance lip syncKling lip syncVeo lip syncRunway lip sync