# Grok Imagine Video 1.5 Text to Video > Generate cinematic videos up to 1080p from a text prompt, with natively synchronized audio in one pass. ## Overview - **Endpoint**: `https://api.segmind.com/v1/grok-imagine-video-1.5-text-to-video` - **Model ID**: `grok-imagine-video-1.5-text-to-video` - **Category**: Text-to-Video Generation - **Type**: Synchronous (Direct response) - **Average Latency**: ~98.0s (30-day average) - **Average Cost**: $3.645 per run (observed across past runs, not a price — see Pricing) - **Provider**: XAI (Grok) ## Pricing | charge | rate | | --- | --- | | 480p video output | $0.108 per second of output video | | 720p video output | $0.189 per second of output video | | 1080p video output | $0.3375 per second of output video | ## Pricing Notes Billed at xAI's exact reported cost plus Segmind's standard margin: $0.10/s at 480p, $0.175/s at 720p, $0.3125/s at 1080p, per second of output video. ## API Information This model uses a **synchronous response pattern**: 1. Make a POST request with your parameters 2. Receive the output directly in the response (binary for images/videos/audio, JSON for text) 3. No polling required - response is immediate ### Input Schema The API accepts the following input parameters: - **`prompt`** (`string`, _required_): Prompt Describe the scene, subject, camera move and sound design. - Default: `"A cat surfing a wave at sunset, cinematic slow motion, the camera tracking alongside as the water roars."` - **`duration`** (`integer`, _optional_): Duration (s) Video length, 1-15 seconds. Billed per second of output video. - Default: `6` - Range: 1 to 15 - **`resolution`** (`string`, _optional_): Resolution Output resolution; higher tiers cost more per second. 480p, 720p or 1080p. - Default: `"480p"` - Options: "480p" (480p), "720p" (720p), "1080p" (1080p) - **`aspect_ratio`** (`string`, _optional_): Aspect Ratio Output aspect ratio. 16:9 landscape, 9:16 social verticals, 1:1 square. - Default: `"16:9"` - Options: "16:9" (16:9), "9:16" (9:16), "1:1" (1:1), "4:3" (4:3), "3:4" (3:4), "3:2" (3:2), "2:3" (2:3) **Required Parameters Example**: ```json { "prompt": "A vintage red sports car speeding along a coastal cliffside highway at golden-hour sunset, cinematic tracking shot alongside the car, the engine roaring as waves crash against the rocks below and wind rushes past." } ``` **Full Example**: ```json { "prompt": "A vintage red sports car speeding along a coastal cliffside highway at golden-hour sunset, cinematic tracking shot alongside the car, the engine roaring as waves crash against the rocks below and wind rushes past.", "duration": 6, "resolution": "720p", "aspect_ratio": "16:9" } ``` ### Output Schema The API returns a synchronous response based on the model type: **For Image/Video/Audio Models**: - Response contains binary data (image/png, video/mp4, audio/mp3) - Content-Type header indicates the media type - Save the response body directly to a file **For Text Models**: - Response is JSON with the generated text - Structure varies by model **HTTP Response Codes**: - **200 - OK**: Request successful, output in response body - **400 - Bad Request**: Invalid parameters - **401 - Unauthorized**: Invalid or missing API key - **404 - Not Found**: Model not found - **406 - Not Acceptable**: Insufficient credits - **429 - Too Many Requests**: Rate limit exceeded - **500 - Server Error**: Internal server error ## About ### Grok Imagine Video 1.5 Text to Video Grok Imagine Video 1.5 Text to Video is xAI's text-to-video model that turns a written prompt into a cinematic clip with natively generated, synchronized audio — no starting image required. #### What is Grok Imagine Video 1.5 Text to Video? Grok Imagine Video 1.5 Text to Video is the generally available text-to-video mode of xAI's Grok Imagine Video 1.5. You describe a scene, subject, camera move, and sound design in plain language, and the model returns an MP4 with picture and audio produced together in a single pass. It is built on xAI's Aurora engine, which renders each clip frame by frame so motion, lighting, and camera trajectory stay coherent across the shot. Outputs run from 1 to 15 seconds at 480p, 720p, or native 1080p, across seven aspect ratios. #### Key Features - **Native synchronized audio**: dialogue, ambient sound, sound effects, and background music generated in the same pass — no separate audio tool or alignment step. - **Prompt-only generation**: create a new scene from text alone, no reference image needed. - **Up to native 1080p** at 24fps, with 480p and 720p tiers for faster drafts. - **Flexible duration** from 1 to 15 seconds and **seven aspect ratios** (16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3). - **Fast, coherent motion** from the autoregressive Aurora engine. #### Best Use Cases Grok Imagine Video 1.5 Text to Video is strongest for social-native short clips for TikTok, Reels, Stories, and X, where native audio removes post-production overhead. It also suits cinematic teasers and ad creative, product and concept b-roll, and rapid pre-visualization for storyboards and camera tests. In testing, 720p prompts returned clean, artifact-free clips with well-synchronized ambient audio and smooth cinematic tracking, generated in roughly 45 seconds. #### Prompt Tips and Output Quality Front-load the key action: the model renders actions described early in the prompt early in the clip, so lead with the main motion, then add camera direction and sound cues. Because audio is generated natively, name the sounds you want ("engine roaring", "waves crashing"). Use 480p for quick drafts, 720p as a balanced showcase, and 1080p for hero shots. Keep clips to 4-6 seconds for the tightest cinematic results. ## Usage Guide ### How to Use Grok Imagine Video 1.5 Text to Video Grok Imagine Video 1.5 Text to Video turns a single text prompt into a cinematic MP4 clip with natively synchronized audio. This guide covers how to prompt it and which settings to pick for each use case. #### Writing an Effective Prompt Structure your prompt in three parts: the subject and action, the camera move, and the sound design. Front-load the key action — Aurora renders the video frame by frame, so actions described early appear early in the clip, while details buried at the end may arrive too late to show clearly. Because audio is generated in the same pass, spell out the sounds you want, such as "engine roaring", "waves crashing", or "soft ambient wind". Add camera language like "slow push-in", "tracking shot", or "handheld pan" to direct movement. #### Choosing Resolution - **480p**: fastest turnaround, ideal for quick drafts and iterating on prompts. - **720p**: the balanced default for social clips and showcases with crisp detail. - **1080p**: native Full HD for final hero shots where sharpness matters most. #### Choosing Aspect Ratio Use 16:9 for widescreen and cinematic shots, 9:16 for TikTok, Reels, and Stories, 1:1 for square social posts and thumbnails, 4:3 or 3:4 for presentations and portraits, and 3:2 or 2:3 for photography-style framing. #### Choosing Duration Duration ranges from 1 to 15 seconds. Keep clips to 4-6 seconds for the tightest, most coherent cinematic results, and extend toward 10-15 seconds for fuller scenes or narrated moments. #### Recommended Settings Tested and validated defaults for a strong first result: - **prompt**: a scene with subject, camera move, and sound cues (e.g. a cinematic tracking shot with engine and surf audio) - **duration**: 6 - **resolution**: 720p - **aspect_ratio**: 16:9 These settings produced clean, artifact-free clips with well-synchronized stereo audio and smooth motion in about 45 seconds. ## FAQ ### Does Grok Imagine Video 1.5 Text to Video generate audio? Yes. Every clip includes native, synchronized audio — dialogue, ambient sound, effects, and music — created in the same generation pass as the video. ### Do I need a starting image? No. This is text-to-video: it generates a full scene from a prompt alone. ### What resolutions and durations are supported? 480p, 720p, and native 1080p, at 1 to 15 seconds per clip. ### What aspect ratios can I use? Seven presets: 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, and 2:3. ### How do I get longer sequences? Generate multiple clips and chain them, continuing from the final frame to maintain motion and lighting continuity. ## Usage Examples ### cURL ```bash curl -X POST "https://api.segmind.com/v1/grok-imagine-video-1.5-text-to-video" \ -H "x-api-key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "A vintage red sports car speeding along a coastal cliffside highway at golden-hour sunset, cinematic tracking shot alongside the car, the engine roaring as waves crash against the rocks below and wind rushes past.", "duration": 6, "resolution": "720p", "aspect_ratio": "16:9" }' ``` ### Python ```python import requests import json api_key = "YOUR_API_KEY" url = "https://api.segmind.com/v1/grok-imagine-video-1.5-text-to-video" data = { "prompt": "A vintage red sports car speeding along a coastal cliffside highway at golden-hour sunset, cinematic tracking shot alongside the car, the engine roaring as waves crash against the rocks below and wind rushes past.", "duration": 6, "resolution": "720p", "aspect_ratio": "16:9" } response = requests.post( url, json=data, headers={ 'x-api-key': api_key, 'Content-Type': 'application/json' } ) if response.status_code == 200: # For image/video/audio models, response.content contains the binary data with open('output.png', 'wb') as f: f.write(response.content) print('Generation complete, saved to output.png') else: print(f"Error: {response.status_code}") print(response.text) ``` ### JavaScript ```javascript const apiKey = 'YOUR_API_KEY'; const url = 'https://api.segmind.com/v1/grok-imagine-video-1.5-text-to-video'; const data = { "prompt": "A vintage red sports car speeding along a coastal cliffside highway at golden-hour sunset, cinematic tracking shot alongside the car, the engine roaring as waves crash against the rocks below and wind rushes past.", "duration": 6, "resolution": "720p", "aspect_ratio": "16:9" }; const response = await fetch(url, { method: 'POST', headers: { 'x-api-key': apiKey, 'Content-Type': 'application/json', }, body: JSON.stringify(data), }); if (response.ok) { // For image/video/audio models, response contains binary data const blob = await response.blob(); const downloadUrl = URL.createObjectURL(blob); // Create download link const a = document.createElement('a'); a.href = downloadUrl; a.download = 'output.png'; a.click(); console.log('Generation complete'); } ``` ## Additional Resources ### Documentation - [Model Playground](https://www.segmind.com/models/grok-imagine-video-1.5-text-to-video) - [API Documentation](https://www.segmind.com/models/grok-imagine-video-1.5-text-to-video/api) - [Pricing Details](https://www.segmind.com/models/grok-imagine-video-1.5-text-to-video/pricing) - [Platform Documentation](https://docs.segmind.com/)