# Kling O3 Text-to-Video > Generate cinematic AI videos up to 15 seconds with native audio, multi-shot control, and physics-accurate motion via API. ## Overview - **Endpoint**: `https://api.segmind.com/v1/kling-o3-text2video` - **Model ID**: `kling-o3-text2video` - **Category**: Text-to-Video Generation - **Type**: Synchronous (Direct response) - **Average Latency**: ~180.8s (30-day average) - **Average Cost**: $1.96 per run (observed across past runs, not a price — see Pricing) - **Provider**: EACHLABS ## Pricing | duration | mode | generate_audio | cost | | --- | --- | --- | --- | | 3 | std | false | 0.63 | | 3 | std | true | 0.84 | | 3 | pro | false | 0.84 | | 3 | pro | true | 1.05 | | 4 | std | false | 0.84 | | 4 | std | true | 1.12 | | 4 | pro | false | 1.12 | | 4 | pro | true | 1.4 | | 5 | std | false | 1.05 | | 5 | std | true | 1.4 | | 5 | pro | false | 1.4 | | 5 | pro | true | 1.75 | | 6 | std | false | 1.26 | | 6 | std | true | 1.68 | | 6 | pro | false | 1.68 | | 6 | pro | true | 2.1 | | 7 | std | false | 1.47 | | 7 | std | true | 1.96 | | 7 | pro | false | 1.96 | | 7 | pro | true | 2.45 | | 8 | std | false | 1.68 | | 8 | std | true | 2.24 | | 8 | pro | false | 2.24 | | 8 | pro | true | 2.8 | | 9 | std | false | 1.89 | | 9 | std | true | 2.52 | | 9 | pro | false | 2.52 | | 9 | pro | true | 3.15 | | 10 | std | false | 2.1 | | 10 | std | true | 2.8 | | 10 | pro | false | 2.8 | | 10 | pro | true | 3.5 | | 11 | std | false | 2.31 | | 11 | std | true | 3.08 | | 11 | pro | false | 3.08 | | 11 | pro | true | 3.85 | | 12 | std | false | 2.52 | | 12 | std | true | 3.36 | | 12 | pro | false | 3.36 | | 12 | pro | true | 4.2 | | 13 | std | false | 2.73 | | 13 | std | true | 3.64 | | 13 | pro | false | 3.64 | | 13 | pro | true | 4.55 | | 14 | std | false | 2.94 | | 14 | std | true | 3.92 | | 14 | pro | false | 3.92 | | 14 | pro | true | 4.9 | | 15 | std | false | 3.15 | | 15 | std | true | 4.2 | | 15 | pro | false | 4.2 | | 15 | pro | true | 5.25 | ## API Information This model uses a **synchronous response pattern**: 1. Make a POST request with your parameters 2. Receive the output directly in the response (binary for images/videos/audio, JSON for text) 3. No polling required - response is immediate ### Input Schema The API accepts the following input parameters: - **`mode`** (`string`, _optional_): Mode Select quality mode. Standard for faster generation, Pro for higher quality. - Default: `"pro"` - Options: "std" (Standard), "pro" (Pro) - **`prompt`** (`string`, _required_): Prompt Describe the video content. - **`duration`** (`string`, _optional_): Duration (seconds) Length of the output video in seconds. - Default: `"5"` - Options: "3" (3), "4" (4), "5" (5), "6" (6), "7" (7), "8" (8), "9" (9), "10" (10), "11" (11), "12" (12), "13" (13), "14" (14), "15" (15) - **`cfg_scale`** (`number`, _optional_): CFG Scale Prompt adherence strength. - Default: `0.5` - Range: 0 to 1 - **`shot_type`** (`string`, _optional_): Shot Type Camera shot type for the video. - Default: `"customize"` - **`voice_ids`** (`array`, _optional_): Voice IDs Voice IDs for audio. - Item type: string - **`aspect_ratio`** (`string`, _optional_): Aspect Ratio Set the aspect ratio for the output video. - Default: `"16:9"` - Options: "16:9" (16:9), "9:16" (9:16), "1:1" (1:1) - **`multi_prompt`** (`array`, _optional_): Multi Prompt Array of prompts with durations for multi-segment video. - Item type: object - **`generate_audio`** (`boolean`, _optional_): Generate Audio Enable to generate synchronized audio with the video. - Default: `false` - **`negative_prompt`** (`string`, _optional_): Negative Prompt Specify elements to avoid. - Default: `"blur, distort, and low quality"` **Required Parameters Example**: ```json { "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, breathtaking natural scenery, ultra detailed" } ``` **Full Example**: ```json { "mode": "pro", "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, breathtaking natural scenery, ultra detailed", "duration": "5", "cfg_scale": 0.5, "shot_type": "customize", "voice_ids": [], "aspect_ratio": "16:9", "multi_prompt": [], "generate_audio": false, "negative_prompt": "blur, distort, and low quality" } ``` ### Output Schema The API returns a synchronous response based on the model type: **For Image/Video/Audio Models**: - Response contains binary data (image/png, video/mp4, audio/mp3) - Content-Type header indicates the media type - Save the response body directly to a file **For Text Models**: - Response is JSON with the generated text - Structure varies by model **HTTP Response Codes**: - **200 - OK**: Request successful, output in response body - **400 - Bad Request**: Invalid parameters - **401 - Unauthorized**: Invalid or missing API key - **404 - Not Found**: Model not found - **406 - Not Acceptable**: Insufficient credits - **429 - Too Many Requests**: Rate limit exceeded - **500 - Server Error**: Internal server error ## About ### Kling O3 Text-to-Video — AI Video Generation Model #### What is Kling O3 Text-to-Video? Kling O3 (Video 3.0 Omni) is the flagship text-to-video model from Kuaishou Technology, launched in February 2026. It transforms natural language descriptions into high-quality video clips up to 15 seconds long, with support for native synchronized audio generation in a single inference pass. Built on the "Omni One" architecture, Kling O3 combines 3D Spacetime Joint Attention with Chain-of-Thought reasoning — enabling the model to think through scene composition, motion physics, and camera behavior before rendering a single frame. The result is cinematic video that respects real-world physics and maintains consistent characters across shots. Unlike earlier text-to-video systems that treated generation as a simple mapping task, Kling O3 operates more like a film director: it interprets your prompt for narrative intent, decomposes it into scene elements, plans motion paths and lighting, and executes the full sequence — all in one generation. #### Key Features - **Native audio generation**: Dialogue, ambient sound, and music are synthesized simultaneously with the video — no post-processing required. Lip-sync, spatial audio, and breath timing emerge naturally from the same generation pass. - **Multi-shot storyboarding**: Generate up to 6 distinct camera shots with individual prompts and durations in a single API call, enabling full scene narratives without manual stitching. - **Physics-accurate motion**: Built-in physics engine models gravity, balance, deformation, collision, and inertia for realistic character and object movement. - **Flexible aspect ratios**: 16:9 for widescreen, 9:16 for mobile/social, and 1:1 for square content. - **Standard and Pro quality modes**: Standard delivers faster turnaround for drafts and iteration; Pro produces cinematic-grade output for final delivery. - **Extended duration**: 3 to 15 second clip lengths — significantly longer than most competing models. #### Best Use Cases **Content creation and marketing**: Generate product showcase videos, brand storytelling clips, and social media content at scale without a full production team. The character consistency features are especially useful for series content requiring the same subject across multiple videos. **Film pre-visualization**: Directors and storyboard artists can quickly prototype shot sequences, camera movements, and scene compositions before committing to production shoots. **E-commerce**: Demonstrate products in context — lifestyle scenes, environments, and dynamic product shots generated from a brief description. **Education and training**: Create explainer videos with synchronized narration, visualize abstract concepts, or produce multilingual training scenarios using the built-in multi-language audio support (English, Chinese, Japanese, Korean, Spanish). **Developers and product teams**: Integrate Kling O3 into content pipelines via the Segmind API endpoint at `https://api.segmind.com/v1/kling-o3-text2video`. Average generation time is approximately 75 seconds per clip. #### Prompt Tips and Output Quality Write prompts as scene directions, not image descriptions. A strong Kling O3 prompt specifies the **subject**, **action**, **camera movement**, **lighting**, and **mood** — for example: "Medium shot of a woman walking through a neon-lit Tokyo street at night, camera panning slowly, warm glow reflecting on wet pavement, cinematic and moody." Use the `cfg_scale` parameter to control how closely the output follows your prompt: values around 0.5 are a good default; push toward 1.0 for strict adherence. Use `negative_prompt` to suppress common artifacts — including `blur, distort, low quality, shaky camera` covers most cases. For Pro mode, generation produces sharper motion and better lighting consistency, especially for longer clips above 8 seconds. ## Usage Guide ### How to Use Kling O3 Text-to-Video Kling O3 is Kuaishou's most capable text-to-video model, designed for creators and developers who need cinematic, physics-accurate video from natural language prompts. Here's how to get the best results. #### Writing Effective Prompts Think like a director, not a photographer. Kling O3 responds best to directorial language that describes action, movement, and atmosphere — not just what a scene looks like. **Good prompt structure**: `[Shot type] of [subject] [action] in [setting], [camera movement], [lighting], [mood/style].` Example: "Close-up of a chef plating a dish in a modern kitchen, slow push-in, warm overhead lighting, cinematic and precise." Avoid vague keywords like "beautiful" or "cool." Describe what the camera should see and how it should move. #### Choosing Mode and Duration Use **Standard mode** for fast iteration and testing your prompt. Once you're happy with the composition, switch to **Pro mode** for final output — it produces sharper motion, better lighting consistency, and higher visual fidelity, especially for clips over 5 seconds. For **duration**, start with 5 seconds to validate your prompt, then extend. Clips over 10 seconds require more detailed prompts to maintain coherence across the full length. #### Using Multi-Prompt for Scene Sequences Pass an array to `multi_prompt` to generate multi-shot sequences in a single call. Each segment takes a `prompt` and `duration`: ```json [ {"prompt": "Aerial view of a coastline at sunrise, camera drifting forward", "duration": "5"}, {"prompt": "Medium shot of a surfer paddling out, golden light, slow motion", "duration": "5"}, {"prompt": "Close-up of a wave crashing, water detail, cinematic", "duration": "5"} ] ``` This is ideal for social media reels, explainer videos, and narrative sequences. #### Audio Generation Enable `generate_audio: true` for scenes with speech, environmental sounds, or music. Pair with `voice_ids` from the Kling Create Voice endpoint to use specific cloned or preset voices — reference them in your prompt as `<<>>`. #### Suppressing Artifacts Use `negative_prompt` to reduce common quality issues. A reliable default: `blur, distort, low quality, shaky camera, watermark`. For scenes with detailed hands or text on screen, add `distorted hands, garbled text` to further improve output consistency. #### CFG Scale Keep `cfg_scale` around 0.5 for most use cases. Increase toward 0.8-1.0 when your prompt is very specific and you want strict adherence. Lower values (0.2-0.4) are useful for experimental or artistic outputs where creative variation is welcome. --- ## FAQ ### What is the maximum video length Kling O3 can generate? Kling O3 supports clips from 3 to 15 seconds in a single generation. ### Does Kling O3 generate audio automatically? Audio generation is optional. Enable the `generate_audio` parameter to get synchronized dialogue, ambient sound, and music alongside the video. You can also pass custom voice IDs via the `voice_ids` parameter. ### What is the difference between Standard and Pro mode? Standard mode is faster and costs less — ideal for drafts and iteration. Pro mode produces higher-quality cinematic output with better lighting, motion, and detail, particularly noticeable in clips over 5 seconds. ### Can I generate multi-shot videos with different scenes? Yes. Use the `multi_prompt` parameter to pass an array of scene segments, each with its own prompt and duration. This supports up to 6 camera cuts in a single generation. ### What aspect ratios are supported? 16:9 (landscape), 9:16 (portrait/mobile), and 1:1 (square). Select based on your target platform. ### How long does generation take? Average latency is approximately 75 seconds. Longer durations and Pro mode will take slightly more time. --- ## Usage Examples ### cURL ```bash curl -X POST "https://api.segmind.com/v1/kling-o3-text2video" \ -H "x-api-key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "mode": "pro", "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, breathtaking natural scenery, ultra detailed", "duration": "5", "cfg_scale": 0.5, "shot_type": "customize", "voice_ids": [], "aspect_ratio": "16:9", "multi_prompt": [], "generate_audio": false, "negative_prompt": "blur, distort, and low quality" }' ``` ### Python ```python import requests import json api_key = "YOUR_API_KEY" url = "https://api.segmind.com/v1/kling-o3-text2video" data = { "mode": "pro", "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, breathtaking natural scenery, ultra detailed", "duration": "5", "cfg_scale": 0.5, "shot_type": "customize", "voice_ids": [], "aspect_ratio": "16:9", "multi_prompt": [], "generate_audio": false, "negative_prompt": "blur, distort, and low quality" } response = requests.post( url, json=data, headers={ 'x-api-key': api_key, 'Content-Type': 'application/json' } ) if response.status_code == 200: # For image/video/audio models, response.content contains the binary data with open('output.png', 'wb') as f: f.write(response.content) print('Generation complete, saved to output.png') else: print(f"Error: {response.status_code}") print(response.text) ``` ### JavaScript ```javascript const apiKey = 'YOUR_API_KEY'; const url = 'https://api.segmind.com/v1/kling-o3-text2video'; const data = { "mode": "pro", "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, breathtaking natural scenery, ultra detailed", "duration": "5", "cfg_scale": 0.5, "shot_type": "customize", "voice_ids": [], "aspect_ratio": "16:9", "multi_prompt": [], "generate_audio": false, "negative_prompt": "blur, distort, and low quality" }; const response = await fetch(url, { method: 'POST', headers: { 'x-api-key': apiKey, 'Content-Type': 'application/json', }, body: JSON.stringify(data), }); if (response.ok) { // For image/video/audio models, response contains binary data const blob = await response.blob(); const downloadUrl = URL.createObjectURL(blob); // Create download link const a = document.createElement('a'); a.href = downloadUrl; a.download = 'output.png'; a.click(); console.log('Generation complete'); } ``` ## Additional Resources ### Documentation - [Model Playground](https://www.segmind.com/models/kling-o3-text2video) - [API Documentation](https://www.segmind.com/models/kling-o3-text2video/api) - [Pricing Details](https://www.segmind.com/models/kling-o3-text2video/pricing) - [Platform Documentation](https://docs.segmind.com/)