MiniMax Hailuo H3 Reference to Video

Keep characters and products consistent in 2K reference-to-video.

Example output

MiniMax Hailuo H3 Reference to Video — Image-to-Video Generation

What is MiniMax Hailuo H3 Reference to Video?

MiniMax Hailuo H3 Reference to Video is the reference-driven mode of MiniMax H3 (also called Hailuo 3.0), MiniMax's omni-modal video model. You supply one to five reference images and a text prompt, and the model generates a native 2K clip that keeps your subject — a character, face, or product — consistent from the first frame to the last. It is built for creators, developers, and product teams who need identity to hold across shots instead of drifting between generations.

Under the hood, H3 reads images and text as a single context, so the reference images pin appearance while the prompt drives the action, camera, and lighting. Clips run 4 to 15 seconds at a film-standard 24fps.

Key Features

  • Reference-driven consistency: pin a subject with 1-5 images so faces, styling, and products stay identical across the clip.
  • Native 2K output: high-resolution frames at 24fps for a timeline-ready look.
  • Flexible duration: 4-15 seconds in one generation, controlled by a single parameter.
  • Aspect ratio control: adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 for landscape, square, or vertical delivery.
  • Prompt-led motion: describe scene, subject, and camera in natural language up to 7000 characters.

Best Use Cases

Reference-to-video shines for product ads with consistent product framing, music videos and short films with a recurring character, anime or illustrated content that holds its linework, and UGC-style social clips. It is a strong fit for advertising, branding, e-commerce, product design, and UI/UX work where a subject must look the same in every shot.

Prompt Tips and Output Quality

Give each reference image one clear job and use clean, well-lit, consistent front views. Name the traits you need held — hair, wardrobe, expression, or product angle — directly in the prompt rather than trusting the image alone. Describe an action that runs the full length of the clip instead of a frozen frame, and specify camera movement and lighting for cinematic results.

FAQs

Is this the same as Hailuo 3.0? Yes. MiniMax H3 is the official name; Hailuo 3.0 is the common nickname for the same model.

Is it the same as MiniMax M3? No. M3 is MiniMax's text and agent model; H3 is the video model.

How many reference images can I use? Between one and five per generation.

What resolution does it output? Native 2K is the available tier on this endpoint.

How long can a clip be? Any whole number of seconds from 4 to 15.

What aspect ratios are supported? Adaptive plus 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.