# Grok Imagine Video 1.5 (Preview) > Turn a still image into cinematic video with natively synchronized audio, using xAI's leaderboard-topping image-to-video model. ## Overview - **Endpoint**: `https://api.segmind.com/v1/grok-imagine-video-1.5-preview` - **Model ID**: `grok-imagine-video-1.5-preview` - **Category**: Image-to-Video Generation - **Type**: Synchronous (Direct response) - **Average Latency**: ~48.0s (30-day average) - **Average Cost**: $1.709 per run (observed across past runs, not a price — see Pricing) - **Provider**: XAI (Grok) ## Pricing | charge | rate | | --- | --- | | 480p video output | $0.108 per second of output video | | 720p video output | $0.189 per second of output video | | Each input or reference image | $0.0135 per image | ## API Information This model uses a **synchronous response pattern**: 1. Make a POST request with your parameters 2. Receive the output directly in the response (binary for images/videos/audio, JSON for text) 3. No polling required - response is immediate ### Input Schema The API accepts the following input parameters: - **`prompt`** (`string`, _required_): Prompt Describe motion, camera moves and sound. Keep it short and motion-focused. - Default: `"A cat surfing a wave at sunset, cinematic slow motion"` - **`image`** (`File (URL)`, _optional_): Image Image to animate, URL or base64. Match aspect ratio to image orientation. - Default: `"https://segmind-resources.s3.amazonaws.com/input/grok-imagine-video-1.5-preview-input-cat-surf.jpeg"` - **`duration`** (`integer`, _optional_): Duration (s) Video length, 1-15 seconds, billed per output second. Use 6-10s for most clips. - Default: `6` - Range: 1 to 15 - **`resolution`** (`string`, _optional_): Resolution Output resolution, 480p or 720p; higher costs more. Use 720p for hero assets. - Default: `"480p"` - Options: "480p" (480p), "720p" (720p) - **`aspect_ratio`** (`string`, _optional_): Aspect Ratio Output aspect ratio. 16:9 landscape, 9:16 social verticals, 1:1 square. - Default: `"16:9"` - Options: "16:9" (16:9), "9:16" (9:16), "1:1" (1:1), "4:3" (4:3), "3:4" (3:4), "3:2" (3:2), "2:3" (2:3) **Required Parameters Example**: ```json { "prompt": "A cat surfing a wave at sunset, cinematic slow motion" } ``` **Full Example**: ```json { "prompt": "A cat surfing a wave at sunset, cinematic slow motion", "image": "https://segmind-resources.s3.amazonaws.com/input/grok-imagine-video-1.5-preview-input-cat-surf.jpeg", "duration": 6, "resolution": "480p", "aspect_ratio": "16:9" } ``` ### Output Schema The API returns a synchronous response based on the model type: **For Image/Video/Audio Models**: - Response contains binary data (image/png, video/mp4, audio/mp3) - Content-Type header indicates the media type - Save the response body directly to a file **For Text Models**: - Response is JSON with the generated text - Structure varies by model **HTTP Response Codes**: - **200 - OK**: Request successful, output in response body - **400 - Bad Request**: Invalid parameters - **401 - Unauthorized**: Invalid or missing API key - **404 - Not Found**: Model not found - **406 - Not Acceptable**: Insufficient credits - **429 - Too Many Requests**: Rate limit exceeded - **500 - Server Error**: Internal server error ## About ### Grok Imagine Video 1.5 (Preview) — Image-to-Video Generation Model Grok Imagine Video 1.5 Preview is xAI's latest image-to-video AI model. It turns a single still image into a fluid, cinematic video clip — with natively synchronized audio — guided by a natural-language prompt. #### What is Grok Imagine Video 1.5? Released in preview on May 30, 2026, Grok Imagine Video 1.5 animates a starting frame into up to 15 seconds of 24fps video at 480p or 720p. Give it an image and a prompt describing the motion, and it renders camera moves, atmosphere, and physics while staying faithful to the detail and lighting of your source image. It debuted at #1 on the Artificial Analysis Image-to-Video Arena leaderboard, ahead of Runway, Kling, and Veo. #### Key Features - **Native synchronized audio** — dialogue, sound effects, ambient sound, and music are generated in the same inference pass, not added afterwards. - **Source-image fidelity** — the output continues your image rather than reinterpreting it, preserving subject, lighting, and composition. - **Promptable camera direction** — describe push-ins, pans, pacing, and sound design in plain language. - **Flexible output** — 1-15 second clips, 480p or 720p, and seven aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3). - **Fast generation** — a 6-second clip typically renders in about 30 seconds. #### Best Use Cases Animate product shots into lifestyle video ads, turn concept art or storyboards into moving sequences, produce vertical 9:16 social clips with sound, and chain shots together — stage each frame as an image, animate it, and cut the clips into longer scenes with a consistent look. #### Prompt Tips and Output Quality The input image anchors the content, so keep prompts short and motion-focused: describe the camera move, the subject's action, and the soundscape. In our testing, outputs tracked the source image closely with coherent, dynamic motion and a clean synchronized audio track. Match the aspect ratio to your input image orientation for best framing. ## Usage Guide ### How to Use Grok Imagine Video 1.5 (Preview) Effectively Grok Imagine Video 1.5 animates a still image into a short cinematic clip with native synchronized audio. Quality depends mostly on two things: a strong input image and a short, motion-focused prompt. #### Recommended Settings (from our testing) These tested values produced a clean, representative result: - **prompt**: A cat surfing a wave at sunset, cinematic slow motion - **duration**: 6 seconds - **resolution**: 480p - **aspect_ratio**: 16:9 A 6-second 480p clip rendered in about 30 seconds with coherent motion, strong fidelity to the input image, and a synchronized ambient audio track. #### Choosing Parameters by Use Case **Social media clips** — use 9:16 with a vertical input image and 6-8 seconds. 480p is fine for feeds. **Hero and marketing assets** — switch to 720p. Cost scales per output second and with resolution, so keep duration tight. **Longer sequences** — generate each shot from its own staged frame and chain clips; the model keeps a consistent look when frames share style and lighting. #### Prompting Tips The image anchors subject, lighting, and composition — the prompt steers motion. Describe three things: the subject's action, the camera move (push-in, pan, tracking), and the sound design. Avoid re-describing what is already visible in the frame. #### Input Image Guidance Use a sharp, well-lit image whose orientation matches your chosen aspect ratio. The model preserves detail from the source frame, so artifacts in the input carry into the video. #### Duration and Cost Control Billing is per second of output video. Start at 6 seconds for iteration, then extend to 10-15 seconds only for final renders. ## FAQ ### Does Grok Imagine Video 1.5 support text-to-video? No. An input image is required. Generate a frame with a text-to-image model first, then animate it. ### Does it generate sound? Yes — audio is generated natively and synchronized with the video, a standout versus most image-to-video models. ### How long can the videos be? 1 to 15 seconds per clip. Chain multiple shots for longer sequences. ### What resolutions are supported? 480p and 720p at 24fps, across seven aspect ratios. ### Can I control the camera? Yes. Describe camera moves like slow push-ins, pans, or tracking shots directly in the prompt. ### How fast is it? A 6-second 480p clip generates in roughly 30 seconds via the Segmind API. ## Usage Examples ### cURL ```bash curl -X POST "https://api.segmind.com/v1/grok-imagine-video-1.5-preview" \ -H "x-api-key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "A cat surfing a wave at sunset, cinematic slow motion", "image": "https://segmind-resources.s3.amazonaws.com/input/grok-imagine-video-1.5-preview-input-cat-surf.jpeg", "duration": 6, "resolution": "480p", "aspect_ratio": "16:9" }' ``` ### Python ```python import requests import json api_key = "YOUR_API_KEY" url = "https://api.segmind.com/v1/grok-imagine-video-1.5-preview" data = { "prompt": "A cat surfing a wave at sunset, cinematic slow motion", "image": "https://segmind-resources.s3.amazonaws.com/input/grok-imagine-video-1.5-preview-input-cat-surf.jpeg", "duration": 6, "resolution": "480p", "aspect_ratio": "16:9" } response = requests.post( url, json=data, headers={ 'x-api-key': api_key, 'Content-Type': 'application/json' } ) if response.status_code == 200: # For image/video/audio models, response.content contains the binary data with open('output.png', 'wb') as f: f.write(response.content) print('Generation complete, saved to output.png') else: print(f"Error: {response.status_code}") print(response.text) ``` ### JavaScript ```javascript const apiKey = 'YOUR_API_KEY'; const url = 'https://api.segmind.com/v1/grok-imagine-video-1.5-preview'; const data = { "prompt": "A cat surfing a wave at sunset, cinematic slow motion", "image": "https://segmind-resources.s3.amazonaws.com/input/grok-imagine-video-1.5-preview-input-cat-surf.jpeg", "duration": 6, "resolution": "480p", "aspect_ratio": "16:9" }; const response = await fetch(url, { method: 'POST', headers: { 'x-api-key': apiKey, 'Content-Type': 'application/json', }, body: JSON.stringify(data), }); if (response.ok) { // For image/video/audio models, response contains binary data const blob = await response.blob(); const downloadUrl = URL.createObjectURL(blob); // Create download link const a = document.createElement('a'); a.href = downloadUrl; a.download = 'output.png'; a.click(); console.log('Generation complete'); } ``` ## Additional Resources ### Documentation - [Model Playground](https://www.segmind.com/models/grok-imagine-video-1.5-preview) - [API Documentation](https://www.segmind.com/models/grok-imagine-video-1.5-preview/api) - [Pricing Details](https://www.segmind.com/models/grok-imagine-video-1.5-preview/pricing) - [Platform Documentation](https://docs.segmind.com/)