Grok Imagine Video 1.5 Reference to Video

Character-consistent video from up to 7 reference images.

Example output

Grok Imagine Video 1.5 Reference to Video — Image-to-Video Generation

What is Grok Imagine Video 1.5 Reference to Video?

Grok Imagine Video 1.5 Reference to Video is xAI's identity-preserving video model. You supply one to seven reference images plus a prompt, and the model carries the people, products, wardrobe, or style from those images through a generated clip. Unlike standard image-to-video, the references guide what appears in the scene without locking the first frame, so you can keep a character and change the action, setting, or camera. Built on xAI's Aurora engine, it renders the picture and native synchronized audio together in a single pass at 480p or 720p.

Key Features

  • Subject consistency from 1-7 reference images: faces, products, clothing, and locations.
  • Native synchronized audio — dialogue, sound effects, and ambient sound generated with the video.
  • 480p and 720p output at 24fps, with 16:9, 9:16, 1:1, and other aspect ratios.
  • Prompt-driven motion, camera moves, and pacing while references hold appearance.
  • References steer the scene without freezing the opening frame.

Best Use Cases

Reference-to-video shines for character-consistent storytelling, virtual try-on, and product placement. In our testing, a single studio portrait of a red-haired explorer stayed on-model — same face, braids, jacket, and satchel — across a full clip of her hiking a misty golden-hour trail. Use it to keep a brand mascot or spokesperson consistent across shots, drop a product into fresh scenes, stage fashion and try-on clips, or produce short social videos where the subject must match every time.

Prompt Tips and Output Quality

Provide clean, well-lit reference images and let the prompt describe the action, camera move, and sound — not the subject, which the references already fix. Naming each reference in the prompt helps bind it to its role in the scene. Short 5-6 second clips produced sharp, coherent motion and stable identity in testing. For social verticals, switch aspect ratio to 9:16; step up to 720p when you need extra detail.

FAQs

  • How many reference images can I use? Up to seven per generation.
  • Does it lock the first frame? No — references steer content without freezing the opening frame.
  • Does it generate audio? Yes, synchronized audio renders together with the video.
  • What resolutions are supported? 480p and 720p; 1080p is not available for reference-to-video.
  • Is a prompt required? Yes — a descriptive prompt directs the motion and scene.
  • How is it different from image-to-video? Image-to-video animates a single starting frame, while reference-to-video uses images as a visual guide for identity, objects, and style.