Models / Hailuo (MiniMax) / Image to video
Hailuo (MiniMax) image to video prompting guide
Last checked 2026-10-11Hailuo / MiniMax H3 · MiniMax H3 (hailuo3: t2v/i2v/r2v/v2v), H3 Max (h3_max: t2v/i2v first+last), Hailuo 2.3
How to prompt Hailuo (MiniMax) in image-to-video (start frame) mode. Every point links its source; community tips are marked. Your agent gets the same guide through the Atlas MCP (get_guide).
Prompt structure
first frame instruction→style from image→next action→camera amplitude speed→dialogue→soundscape
- Three fields: integrated_multimodal_description (visuals, action, shots, dialogue on the timeline), overall_soundscape (1–4 sentences), non_diegetic_music (or N/A). source ↗
What the text should describe
- Open with: 'For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.' Then establish the image's style/subjects/scene and describe what happens next; keep identity, clothing, colours and layout consistent. source ↗
- On fal an image that literally opens the shot goes to Image to Video (image_url + optional end_image_url); an image used for identity/style goes to Reference to Video. source ↗
What not to describe
- Runway hailuo3: first/last keyframe mode and reference-image mode cannot be mixed in one promptImage array. source ↗
Camera vocabulary
- Camera = motion type + amplitude + speed as a sentence: 'The camera pushes in with small amplitude at slow speed toward ...'. Types: zoom, push/pull, pan, truck, tilt, pedestal, arc, tracking, static, shake, POV, roll. source ↗
Responds well to
- On-screen text in double quotes, verbatim. source ↗
Duration and limits
- H3 Reference to Video: up to 9 images, 3 videos (2–15 s), 3 audio, 12 files; Runway hailuo3 768p/2k, h3_max 480p/768p. source ↗
Multi-shot
- [Shot 1] (no timestamp), then '[Shot 2] At 00:03.500, the camera cuts to ...' with strictly increasing times. source ↗
Avoid and failure modes
- Use a cut only for new information; small angle changes = camera motion. source ↗
Negative prompt
Supported via .
Real prompts, credited
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze ...
Next
All Hailuo (MiniMax) modesHailuo (MiniMax) model, prices & shotsFrom $0.1/s · pricesSeedance image to videoKling image to videoVeo image to videoRunway image to video