# Wan 2.7 Text to Video > Generate cinematic 1080P videos from text with audio sync, multi-shot control, and 15-second duration via Wan 2.7. ## Overview - **Endpoint**: `https://api.segmind.com/v1/wan2.7-t2v` - **Model ID**: `wan2.7-t2v` - **Category**: Text-to-Video Generation - **Type**: Synchronous (Direct response) - **Average Latency**: ~199.9s (30-day average) - **Average Cost**: $1.233 per run (observed across past runs, not a price — see Pricing) - **Provider**: DASHSCOPE ## Pricing | resolution | rate | | --- | --- | | 720p | $0.143 per second | | 1080p | $0.2145 per second | ## API Information This model uses a **synchronous response pattern**: 1. Make a POST request with your parameters 2. Receive the output directly in the response (binary for images/videos/audio, JSON for text) 3. No polling required - response is immediate ### Input Schema The API accepts the following input parameters: - **`prompt`** (`string`, _required_): Prompt Describe scene, subjects, motion, and camera style. Max 5000 chars. - Default: `"A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, dramatic lighting, ultra HD, slow motion"` - **`audio_url`** (`File (URL)`, _optional_): Audio URL URL to an audio file (MP3/WAV) to synchronize with the video. - **`negative_prompt`** (`string`, _optional_): Negative Prompt List elements to exclude from the output (e.g., blurry, watermark, distorted faces). - Default: `"blurry, low quality, distorted, watermark"` - **`resolution`** (`string`, _optional_): Resolution Output resolution tier. - Default: `"720p"` - Options: "720p" (720p), "1080p" (1080p) - **`duration`** (`integer`, _optional_): Duration Length of the generated video in seconds (2-15). - Default: `5` - Range: 2 to 15 - **`ratio`** (`string`, _optional_): Aspect Ratio Aspect ratio of the output video. - Default: `"16:9"` - Options: "16:9" (16:9), "9:16" (9:16), "1:1" (1:1), "4:3" (4:3), "3:4" (3:4) - **`seed`** (`integer`, _optional_): Seed Integer seed for reproducible outputs. - Default: `42` **Required Parameters Example**: ```json { "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, dramatic lighting, ultra HD, slow motion" } ``` **Full Example**: ```json { "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, dramatic lighting, ultra HD, slow motion", "audio_url": "https://example.com/image.jpg", "negative_prompt": "blurry, low quality, distorted, watermark", "resolution": "720p", "duration": 5, "ratio": "16:9", "seed": 42 } ``` ### Output Schema The API returns a synchronous response based on the model type: **For Image/Video/Audio Models**: - Response contains binary data (image/png, video/mp4, audio/mp3) - Content-Type header indicates the media type - Save the response body directly to a file **For Text Models**: - Response is JSON with the generated text - Structure varies by model **HTTP Response Codes**: - **200 - OK**: Request successful, output in response body - **400 - Bad Request**: Invalid parameters - **401 - Unauthorized**: Invalid or missing API key - **404 - Not Found**: Model not found - **406 - Not Acceptable**: Insufficient credits - **429 - Too Many Requests**: Rate limit exceeded - **500 - Server Error**: Internal server error ## About ### Wan 2.7 Text to Video — AI Video Generation API #### What is Wan 2.7? Wan 2.7 is Alibaba's most capable text-to-video model, designed for developers, filmmakers, and content creators who need precise control over AI-generated video. Building on the highly regarded Wan 2.x lineage, version 2.7 delivers a substantial leap in visual quality, motion coherence, and audio synchronization. It generates videos up to 15 seconds long at up to 1080P resolution, supporting five aspect ratios and native audio integration — all through a simple API call. Unlike earlier versions that treated audio as an afterthought, Wan 2.7 supports audio-driven generation from the start: supply an audio URL and the model synchronizes character motion and lip movements with the provided track. This makes it particularly powerful for branded spokesperson content, dubbed video workflows, and music-timed visuals. #### Key Features - **Up to 1080P resolution** at 15 seconds duration, with 720P available for faster iteration - **Native audio synchronization** — provide an audio URL to drive lip-sync and motion timing - **Five aspect ratios** — 16:9, 9:16, 1:1, 4:3, and 3:4 for cross-platform publishing - **Improved motion coherence** — characters and objects move with greater physical plausibility and fewer flickering artifacts - **Cinematic visual quality** — skin textures, fabric movement, and lighting gradients reach commercial-grade 1080P standards - **Reproducible outputs via seed** — lock in a seed to regenerate identical results across iterations #### Best Use Cases **Brand and marketing video**: Generate consistent spokesperson or product demo clips with audio sync, ideal for agencies producing high volumes of branded content. **Social media content**: Use 9:16 ratio at 5-10 seconds for TikTok, Instagram Reels, and YouTube Shorts. 16:9 works for YouTube intros, explainers, and pre-roll ads. **Film and narrative production**: Multi-shot storytelling at 15 seconds with cinematic camera descriptions in the prompt — slow dolly, aerial drone, handheld chase — produces broadcast-adjacent results. **Prototyping and storyboarding**: 720P at 5 seconds for fast iteration on creative concepts before committing to 1080P renders. #### Prompt Tips and Output Quality Wan 2.7 rewards structured, descriptive prompts. Include: (1) the subject and scene, (2) camera movement (e.g., slow pan left, aerial drone descending), (3) lighting and mood (golden hour, soft overcast), and (4) motion details (waves crashing, hair blowing in wind). The model handles cinematic, anime, illustrated, and photorealistic styles — specify your intended aesthetic explicitly. Use the negative_prompt field to suppress common artifacts like blurry, distorted face, watermark, text overlay. For audio-synced content, ensure your audio URL is publicly accessible and under 15 seconds. ## Usage Guide ### How to Use Wan 2.7 Text to Video Wan 2.7 is Alibaba's flagship text-to-video model, combining high-resolution output, audio synchronization, and precise motion control. Here is how to get the most out of it. #### Writing Effective Prompts Structure your prompt in layers: start with the subject (A woman in a red dress), add motion (walking along a rainy Parisian street), then specify the camera style (slow tracking shot from the side), and close with atmosphere (soft evening lighting, cinematic 35mm look). The model handles photorealistic, cinematic, anime, and illustrated aesthetics — name your intended style explicitly. For 15-second multi-shot videos, break your prompt into two or three visual beats separated by commas: Opening on a calm mountain lake at dawn, mist rising, slow push in — cut to a hiker reaching the summit, golden hour backlight, handheld camera — closing on a wide aerial pulling back to reveal the full valley. #### Choosing Resolution and Duration Use 720P during creative development and iteration — it is faster and cheaper. Switch to 1080P for final deliverables, client-facing content, and social publishing. For duration, 5 seconds covers most social media formats; go up to 10-15 seconds for narrative or branded content. #### Audio-Synced Generation If you provide an audio_url, make sure the file is publicly hosted (e.g., on S3 or a CDN) and is under 15 seconds. The model works best with clear speech or music at moderate tempo. Lip-sync is functional for conversational speech; very fast speech (above 150 WPM) may reduce accuracy. #### Aspect Ratio by Platform Match your ratio to your publishing platform: 16:9 for YouTube and desktop, 9:16 for TikTok, Reels, and Shorts, 1:1 for Instagram feed posts, and 4:3 for presentations and slides. #### Negative Prompts and Reproducibility Add negative_prompt values like blurry, watermark, text, distorted face, low quality to clean up outputs. Use a fixed seed integer to lock in a visual style and reproduce it across multiple generations with varied prompts. ## FAQ ### What resolutions does Wan 2.7 support? 720P and 1080P. 720P is faster and costs less; 1080P is suited for final deliverables and high-quality publishing. ### Can I generate longer videos? Yes, up to 15 seconds. Set the duration parameter anywhere from 2 to 15 seconds. ### How does audio-driven generation work? Pass a publicly accessible audio file URL in the audio_url parameter. The model synchronizes character movements and lip motion with the audio track during generation. ### What aspect ratios are available? 16:9 (landscape), 9:16 (portrait), 1:1 (square), 4:3, and 3:4. Choose based on your target platform. ### Does the API return a video file directly? Yes. The response is binary video/mp4 data on HTTP 200. No polling required — the call is synchronous. ### How do I get reproducible outputs? Set a fixed integer seed value. The same seed with identical parameters will reproduce the same video. ## Usage Examples ### cURL ```bash curl -X POST "https://api.segmind.com/v1/wan2.7-t2v" \ -H "x-api-key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, dramatic lighting, ultra HD, slow motion", "audio_url": "https://example.com/image.jpg", "negative_prompt": "blurry, low quality, distorted, watermark", "resolution": "720p", "duration": 5, "ratio": "16:9", "seed": 42 }' ``` ### Python ```python import requests import json api_key = "YOUR_API_KEY" url = "https://api.segmind.com/v1/wan2.7-t2v" data = { "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, dramatic lighting, ultra HD, slow motion", "audio_url": "https://example.com/image.jpg", "negative_prompt": "blurry, low quality, distorted, watermark", "resolution": "720p", "duration": 5, "ratio": "16:9", "seed": 42 } response = requests.post( url, json=data, headers={ 'x-api-key': api_key, 'Content-Type': 'application/json' } ) if response.status_code == 200: # For image/video/audio models, response.content contains the binary data with open('output.png', 'wb') as f: f.write(response.content) print('Generation complete, saved to output.png') else: print(f"Error: {response.status_code}") print(response.text) ``` ### JavaScript ```javascript const apiKey = 'YOUR_API_KEY'; const url = 'https://api.segmind.com/v1/wan2.7-t2v'; const data = { "prompt": "A majestic golden eagle soaring over snow-capped mountain peaks at sunrise, cinematic wide angle shot, dramatic lighting, ultra HD, slow motion", "audio_url": "https://example.com/image.jpg", "negative_prompt": "blurry, low quality, distorted, watermark", "resolution": "720p", "duration": 5, "ratio": "16:9", "seed": 42 }; const response = await fetch(url, { method: 'POST', headers: { 'x-api-key': apiKey, 'Content-Type': 'application/json', }, body: JSON.stringify(data), }); if (response.ok) { // For image/video/audio models, response contains binary data const blob = await response.blob(); const downloadUrl = URL.createObjectURL(blob); // Create download link const a = document.createElement('a'); a.href = downloadUrl; a.download = 'output.png'; a.click(); console.log('Generation complete'); } ``` ## Additional Resources ### Documentation - [Model Playground](https://www.segmind.com/models/wan2.7-t2v) - [API Documentation](https://www.segmind.com/models/wan2.7-t2v/api) - [Pricing Details](https://www.segmind.com/models/wan2.7-t2v/pricing) - [Platform Documentation](https://docs.segmind.com/)