# Grok Imagine Video > Generate cinematic text-to-video and image-to-video clips with native synchronized audio from a prompt or still image. ## Overview - **Endpoint**: `https://api.segmind.com/v1/grok-imagine-video` - **Model ID**: `grok-imagine-video` - **Category**: Text-to-Video Generation - **Type**: Synchronous (Direct response) - **Average Latency**: ~46.0s (30-day average) - **Average Cost**: $0.5458 per run (observed across past runs, not a price — see Pricing) - **Provider**: XAI (Grok) ## Pricing | charge | rate | | --- | --- | | 480p video output | $0.0675 per second of output video | | 720p video output | $0.0945 per second of output video | | Each input or reference image | $0.0027 per image | ## API Information This model uses a **synchronous response pattern**: 1. Make a POST request with your parameters 2. Receive the output directly in the response (binary for images/videos/audio, JSON for text) 3. No polling required - response is immediate ### Input Schema The API accepts the following input parameters: - **`prompt`** (`string`, _required_): Prompt Describes subject, motion, camera and style. Front-load the key action and name audio cues. - Default: `"Ocean waves crashing against dramatic coastal cliffs at golden hour, seagulls calling overhead, sea mist drifting in the air, cinematic slow motion"` - **`image`** (`File (URL)`, _optional_): Image Optional source image (URL or base64) for image-to-video. Omit for text-to-video. - **`duration`** (`integer`, _optional_): Duration (s) Clip length in seconds, 1 to 15. Use 5-6s for social, 10-15s for scenes. - Default: `6` - Range: 1 to 15 - **`resolution`** (`string`, _optional_): Resolution Output resolution: 480p, 720p or 1080p. Use 480p for drafts, 720p/1080p for hero shots. - Default: `"480p"` - Options: "480p" (480p), "720p" (720p), "1080p" (1080p) - **`aspect_ratio`** (`string`, _optional_): Aspect Ratio Frame shape: 16:9 cinematic, 9:16 vertical or social, 1:1 square. - Default: `"16:9"` - Options: "16:9" (16:9), "9:16" (9:16), "1:1" (1:1), "4:3" (4:3), "3:4" (3:4), "3:2" (3:2), "2:3" (2:3) **Required Parameters Example**: ```json { "prompt": "Ocean waves crashing against dramatic coastal cliffs at golden hour, seagulls calling overhead, sea mist drifting in the air, cinematic slow motion" } ``` **Full Example**: ```json { "prompt": "Ocean waves crashing against dramatic coastal cliffs at golden hour, seagulls calling overhead, sea mist drifting in the air, cinematic slow motion", "image": "https://example.com/image.jpg", "duration": 6, "resolution": "480p", "aspect_ratio": "16:9" } ``` ### Output Schema The API returns a synchronous response based on the model type: **For Image/Video/Audio Models**: - Response contains binary data (image/png, video/mp4, audio/mp3) - Content-Type header indicates the media type - Save the response body directly to a file **For Text Models**: - Response is JSON with the generated text - Structure varies by model **HTTP Response Codes**: - **200 - OK**: Request successful, output in response body - **400 - Bad Request**: Invalid parameters - **401 - Unauthorized**: Invalid or missing API key - **404 - Not Found**: Model not found - **406 - Not Acceptable**: Insufficient credits - **429 - Too Many Requests**: Rate limit exceeded - **500 - Server Error**: Internal server error ## About ### Grok Imagine Video: Text-to-Video and Image-to-Video AI Model #### What is Grok Imagine Video? Grok Imagine Video is xAI's video-audio generative model that turns a text prompt or a still image into short, cinematic clips with native synchronized audio. Unlike silent video generators, it produces dialogue, ambient sound, sound effects, and background music in the same generation pass, so the first output is already a coherent audiovisual draft. It supports both text-to-video, where you describe a scene from scratch, and image-to-video, where a source image becomes the starting frame and your prompt drives the motion. Built on xAI's Aurora autoregressive engine, the model renders each frame sequentially from the first frame forward, which keeps subject position, lighting, and camera trajectory stable across the clip. You can generate clips from 1 to 15 seconds at 480p, 720p, or 1080p, in landscape, vertical, square, and other platform-ready aspect ratios. #### Key Features - Native synchronized audio: dialogue with lip-sync, ambient sound, effects, and music generated alongside the video. - Text-to-video and image-to-video from a single endpoint. - Cinematic motion understanding with realistic object interactions and camera moves. - Strong instruction following for controlling subject, action, and style. - Configurable duration (1 to 15 seconds), resolution (480p to 1080p), and aspect ratio. #### Best Use Cases Grok Imagine Video is ideal for social-native short clips for Reels, Shorts, and TikTok, where native audio removes post-production overhead. It excels at animating product shots, portraits, and concept frames from a single still image, and at cinematic teaser generation from reference images. Marketers use it for fast product demos and promo clips, while creators and game designers use it for rapid concept testing and storyboarding before committing to a longer production pipeline. In testing, a text-to-video prompt of crashing ocean waves at golden hour produced a clean 480p clip with clearly audible waves and seagull calls. #### Prompt Tips and Output Quality Write scene-first prompts that name the subject, motion, camera movement, atmosphere, and audio together. Front-load the key action, since the model renders early-described actions early in the clip and may miss details buried at the end. Add explicit audio cues such as waves crashing or birds calling to get richer synchronized sound. For image-to-video, describe only the motion and let the source image anchor identity and composition. ## Usage Guide ### How to Use Grok Imagine Video Grok Imagine Video generates short cinematic clips with native synchronized audio from either a text prompt or a still image. This guide covers how to get clean, controllable results. #### Text-to-Video Set only the `prompt` to generate a clip from scratch. Write a scene-first description that names the subject, the motion, the camera movement, the lighting, and the audio you want. Front-load the most important action because the model renders early-described actions first. Add explicit audio cues such as waves crashing or a crowd cheering to strengthen the synchronized sound. #### Image-to-Video Add a public image URL or base64 string to the `image` field. The source image becomes the first frame and anchors identity, composition, and lighting, so your prompt should describe only the motion, camera move, and atmosphere. This mode is best for animating product shots, portraits, and concept art while preserving the original look. #### Choosing Parameters - `duration`: 1 to 15 seconds. Use 5-6 seconds for social clips and quick iteration; 10-15 seconds for fuller scenes. - `resolution`: 480p for fast drafts and previews; 720p or 1080p for hero shots and final delivery. - `aspect_ratio`: 16:9 for cinematic and landscape, 9:16 for vertical and social, 1:1 for square feeds. #### Recommended Settings These tested text-to-video values produced a clean, artifact-free result with clearly audible native audio: - `prompt`: Ocean waves crashing against dramatic coastal cliffs at golden hour, seagulls calling overhead, sea mist drifting in the air, cinematic slow motion - `duration`: 6 - `resolution`: 480p - `aspect_ratio`: 16:9 Start at 480p and 6 seconds to dial in your prompt cheaply, then raise resolution and duration once the motion and audio look right. For image-to-video, keep the prompt focused on motion and re-run with small prompt tweaks to explore different directions from the same hero frame. ## FAQ ### Does Grok Imagine Video generate audio? Yes. It produces synchronized dialogue, ambient sound, sound effects, and music in the same pass as the video. ### Can it do both text-to-video and image-to-video? Yes. Provide a prompt alone for text-to-video, or add an image to animate a still. ### What is the maximum clip length? Up to 15 seconds per generation; chain clips for longer sequences. ### Which resolutions are supported? 480p, 720p, and 1080p, across landscape, vertical, and square aspect ratios. ### How do I do image-to-video? Pass a public image URL or base64 in the image field and describe the motion you want. ## Usage Examples ### cURL ```bash curl -X POST "https://api.segmind.com/v1/grok-imagine-video" \ -H "x-api-key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "Ocean waves crashing against dramatic coastal cliffs at golden hour, seagulls calling overhead, sea mist drifting in the air, cinematic slow motion", "image": "https://example.com/image.jpg", "duration": 6, "resolution": "480p", "aspect_ratio": "16:9" }' ``` ### Python ```python import requests import json api_key = "YOUR_API_KEY" url = "https://api.segmind.com/v1/grok-imagine-video" data = { "prompt": "Ocean waves crashing against dramatic coastal cliffs at golden hour, seagulls calling overhead, sea mist drifting in the air, cinematic slow motion", "image": "https://example.com/image.jpg", "duration": 6, "resolution": "480p", "aspect_ratio": "16:9" } response = requests.post( url, json=data, headers={ 'x-api-key': api_key, 'Content-Type': 'application/json' } ) if response.status_code == 200: # For image/video/audio models, response.content contains the binary data with open('output.png', 'wb') as f: f.write(response.content) print('Generation complete, saved to output.png') else: print(f"Error: {response.status_code}") print(response.text) ``` ### JavaScript ```javascript const apiKey = 'YOUR_API_KEY'; const url = 'https://api.segmind.com/v1/grok-imagine-video'; const data = { "prompt": "Ocean waves crashing against dramatic coastal cliffs at golden hour, seagulls calling overhead, sea mist drifting in the air, cinematic slow motion", "image": "https://example.com/image.jpg", "duration": 6, "resolution": "480p", "aspect_ratio": "16:9" }; const response = await fetch(url, { method: 'POST', headers: { 'x-api-key': apiKey, 'Content-Type': 'application/json', }, body: JSON.stringify(data), }); if (response.ok) { // For image/video/audio models, response contains binary data const blob = await response.blob(); const downloadUrl = URL.createObjectURL(blob); // Create download link const a = document.createElement('a'); a.href = downloadUrl; a.download = 'output.png'; a.click(); console.log('Generation complete'); } ``` ## Additional Resources ### Documentation - [Model Playground](https://www.segmind.com/models/grok-imagine-video) - [API Documentation](https://www.segmind.com/models/grok-imagine-video/api) - [Pricing Details](https://www.segmind.com/models/grok-imagine-video/pricing) - [Platform Documentation](https://docs.segmind.com/)