# HappyHorse 1.1 > Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or reference images. ## Overview - **Endpoint**: `https://api.segmind.com/v1/happyhorse-1.1` - **Model ID**: `happyhorse-1.1` - **Category**: Image-to-Video Generation - **Type**: Synchronous (Direct response) - **Average Latency**: ~130.1s (30-day average) - **Average Cost**: $2.694 per run (observed across past runs, not a price — see Pricing) - **Provider**: DASHSCOPE ## Pricing | resolution | rate | | --- | --- | | 720p | $0.2002 per second | | 1080p | $0.2574 per second | ## Pricing Notes ### How Pricing Works HappyHorse 1.1 uses per-second, resolution-based pricing. You pay based on the duration and resolution of the generated video. Shorter clips cost less, so you only pay for the output you need. #### Pricing Table | Resolution | Rate | 3s Video | 5s Video | 10s Video | 15s Video | |---|---|---|---|---|---| | 720P | $0.2002/s | $0.60 | $1.00 | $2.00 | $3.00 | | 1080P | $0.2574/s | $0.77 | $1.29 | $2.57 | $3.86 | #### Cost Optimization Tips - Use 720P for iteration and prompt testing, then switch to 1080P only for final renders. 720P costs 22% less per second. - Keep preview clips short (3-5 seconds) during prompt iteration. A 3-second 720P preview costs just $0.60. - Prompt extension is free and significantly improves output quality on short prompts. Always enable it unless you need precise control. - Audio-driven lip-sync and image-to-video modes cost the same as text-to-video at the same resolution and duration. No surcharge for advanced input types. #### Example Scenario A marketing agency producing 20 social video ads per week: 15 preview renders at 720P, 5 seconds each ($1.00 x 15 = $15.00) plus 5 final 1080P renders at 5 seconds ($1.29 x 5 = $6.45). Weekly total: $21.45. Monthly: approximately $85.80 for 80 production-ready videos. Longer 10-second hero videos at 1080P cost $2.57 each. ## API Information This model uses a **synchronous response pattern**: 1. Make a POST request with your parameters 2. Receive the output directly in the response (binary for images/videos/audio, JSON for text) 3. No polling required - response is immediate ### Input Schema The API accepts the following input parameters: - **`prompt`** (`string`, _required_): Prompt Describes the scene, subjects, motion, camera, and audio cues. Use cinematic detail (lighting, lens, dialogue) for richer video and synced native audio. - Default: `"A majestic white horse galloping through golden wheat fields at sunset, cinematic slow motion, warm golden hour lighting, dust particles catching sunlight, shallow depth of field, professional nature documentary style"` - **`image`** (`File (URL)`, _optional_): First Frame (optional) Optional first frame that triggers image-to-video mode. Provide a URL or base64 image to animate it; leave empty for text-to-video. - **`negative_prompt`** (`string`, _optional_): Negative Prompt Lists elements to suppress in the output. Use it to avoid blur, artifacts, watermarks, or text; tune per scene as needed. - Default: `"blurry, low quality, distorted, watermark, text"` - **`resolution`** (`string`, _optional_): Resolution Output resolution. Pick 720P for fast drafts and previews; choose 1080P for final, delivery-ready cinematic video. - Default: `"1080p"` - Options: "720p" (720p), "1080p" (1080p) - **`duration`** (`integer`, _optional_): Duration Video length in seconds, from 3 to 15. Use 3-5s for quick social clips; 8-15s for fuller scenes and narratives. - Default: `5` - Range: 3 to 15 - **`aspect_ratio`** (`string`, _optional_): Aspect Ratio Frame shape for text-to-video (ignored when an image is supplied). Use 16:9 for landscape, 9:16 for vertical, 1:1 for social. - Default: `"16:9"` - Options: "16:9" (16:9), "9:16" (9:16), "1:1" (1:1), "4:3" (4:3), "3:4" (3:4) - **`seed`** (`integer`, _optional_): Seed Reproducibility seed (0-2147483647). Fix a seed to reproduce the same output; leave empty for fresh random variations. - Default: `0` - **`prompt_extend`** (`boolean`, _optional_): Prompt Extend LLM-rewrites short prompts into richer cinematic instructions. Keep on for brief prompts; turn off when you need exact prompt control. - Default: `true` - **`watermark`** (`boolean`, _optional_): Watermark Adds an AI-generated watermark to the output. Disable for clean, production-ready video; enable to label generated content. - Default: `false` - **`reference_images`** (`array`, _optional_): Reference Images Up to 9 reference image URLs or uploads that anchor characters, scenes, style, or products. Use for reference-to-video and multi-scene consistency. - Item type: string **Required Parameters Example**: ```json { "prompt": "A majestic white horse galloping through golden wheat fields at sunset, cinematic slow motion, warm golden hour lighting, dust particles catching sunlight, shallow depth of field, professional nature documentary style" } ``` **Full Example**: ```json { "prompt": "A majestic white horse galloping through golden wheat fields at sunset, cinematic slow motion, warm golden hour lighting, dust particles catching sunlight, shallow depth of field, professional nature documentary style", "image": "https://example.com/image.jpg", "negative_prompt": "blurry, low quality, distorted, watermark, text", "resolution": "1080p", "duration": 5, "aspect_ratio": "16:9", "seed": 0, "prompt_extend": true, "watermark": false, "reference_images": [] } ``` ### Output Schema The API returns a synchronous response based on the model type: **For Image/Video/Audio Models**: - Response contains binary data (image/png, video/mp4, audio/mp3) - Content-Type header indicates the media type - Save the response body directly to a file **For Text Models**: - Response is JSON with the generated text - Structure varies by model **HTTP Response Codes**: - **200 - OK**: Request successful, output in response body - **400 - Bad Request**: Invalid parameters - **401 - Unauthorized**: Invalid or missing API key - **404 - Not Found**: Model not found - **406 - Not Acceptable**: Insufficient credits - **429 - Too Many Requests**: Rate limit exceeded - **500 - Server Error**: Internal server error ## About ### HappyHorse 1.1 — Text & Image to Video with Native Audio #### What is HappyHorse 1.1? HappyHorse 1.1 is Alibaba's unified video-and-audio generation model, built by the Taotian Future Life Lab as the successor to HappyHorse 1.0 — the model that topped the Artificial Analysis Video Arena. Unlike pipelines that bolt dubbing on in post, HappyHorse generates video and synchronized native audio together in a single pass, so dialogue, ambience, and on-screen action line up from the first frame. It also delivers multilingual lip-sync, matching mouth movements to speech across languages like English, Mandarin, Japanese, Korean, German, and French. On Segmind, `happyhorse-1.1` auto-detects three modes from your payload. Send a `prompt` alone for **text-to-video**, add an `image` first frame for **image-to-video**, or pass `reference_images` (up to nine) for **reference-to-video**. Outputs render at 720P or 1080P, in durations from 3 to 15 seconds, across 16:9, 9:16, 1:1, 4:3, and 3:4 aspect ratios. #### Key Features - **Native synchronized audio** — video and audio are jointly generated, no separate dubbing step. - **Multilingual lip-sync** — characters speak with accurate mouth movements across many languages. - **Three auto-detected modes** — text-to-video, image-to-video, and reference-to-video from one endpoint. - **Up to 9 reference images** — anchor characters, scenes, style, and products for multi-scene consistency. - **720P / 1080P** output, **3–15s** duration, and flexible aspect ratios. - **prompt_extend** LLM rewriting, **negative_prompt**, **seed**, and optional **watermark** controls. #### Best Use Cases HappyHorse 1.1 shines for short-form ads and social clips that need character consistency across scenes, global marketing where a single prompt yields multilingual, lip-synced footage, and product or brand series anchored by reference images. Its 1.1 upgrades — improved semantic understanding, cinematic shot control, dynamic motion rendering, stronger subject and visual consistency, richer detail, and more natural character actions and physics — make it well suited to narrative shorts, explainers, music-driven scenes, and storyboard-to-video workflows. #### Prompt Tips and Output Quality Write cinematic prompts: name the subject, the action, the camera move, the lighting, and any audio or dialogue cues. Keep `prompt_extend` on for short prompts to let the model add filmic detail, and turn it off when you need precise control. Use `negative_prompt` to suppress blur, artifacts, and text. For image-to-video, supply a clean first frame; for consistent characters across shots, pass reference images. Fix a `seed` to reproduce a result, and prefer 1080P for final delivery. ## Usage Guide ### How to Use HappyHorse 1.1 HappyHorse 1.1 turns text, a first frame, or reference images into cinematic video with synchronized native audio and multilingual lip-sync. The endpoint auto-detects its mode from the payload, so you control the output entirely through which fields you send. #### Choose your mode - **Text-to-video:** send `prompt` only. Best for original scenes and concepts. Set `aspect_ratio` (16:9 landscape, 9:16 vertical, 1:1 social) to frame the shot. - **Image-to-video:** add an `image` first frame to animate an existing still. `aspect_ratio` is ignored here — the frame defines the shape. - **Reference-to-video:** pass `reference_images` (up to nine) to lock characters, environments, style, or products across a multi-scene project. #### Tune the output - **resolution:** use `720P` for fast drafts and iteration; switch to `1080P` for final, delivery-ready video. - **duration:** 3–15 seconds. Use 3–5s for punchy social clips and 8–15s for fuller narrative beats. - **prompt_extend:** keep `true` for short prompts so the model adds cinematic detail; set `false` when you want exact control over the wording. - **negative_prompt:** suppress blur, artifacts, watermarks, or unwanted text. - **seed:** fix a value to reproduce the same result; leave empty for fresh variations. - **watermark:** keep `false` for clean production output. #### Prompting tips Write like a director: name the subject, the motion, the camera move, the lighting, and any dialogue or ambient audio. Because audio is generated jointly with the video, spelling out speech and sound cues yields better lip-sync and atmosphere. For consistent characters across shots, reuse the same reference images and a fixed seed. Pricing is per second and scales with resolution, so prototype at 720P with short durations, then render your final cut at 1080P once the prompt and references are dialed in. ## FAQ ### Can I upload my own audio for lip-sync? No. HappyHorse 1.1 generates its own native audio; it does not accept an external MP3 or WAV to drive lip-sync. ### How does it pick text-, image-, or reference-to-video? The mode is auto-detected: prompt only is text-to-video, an `image` makes it image-to-video, and `reference_images` makes it reference-to-video. ### How many reference images can I use? Up to nine, to anchor characters, environments, style, and products across scenes. ### What resolutions and durations are supported? 720P or 1080P output, with durations from 3 to 15 seconds. ### Which aspect ratios are available? 16:9, 9:16, 1:1, 4:3, and 3:4 (aspect ratio applies to text-to-video; it is ignored when an image is supplied). ### How is HappyHorse 1.1 different from 1.0? It adds production native audio, multilingual lip-sync, up to nine reference images, 1080P, and improved motion, consistency, and detail. ## Usage Examples ### cURL ```bash curl -X POST "https://api.segmind.com/v1/happyhorse-1.1" \ -H "x-api-key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "A majestic white horse galloping through golden wheat fields at sunset, cinematic slow motion, warm golden hour lighting, dust particles catching sunlight, shallow depth of field, professional nature documentary style", "image": "https://example.com/image.jpg", "negative_prompt": "blurry, low quality, distorted, watermark, text", "resolution": "1080p", "duration": 5, "aspect_ratio": "16:9", "seed": 0, "prompt_extend": true, "watermark": false, "reference_images": [] }' ``` ### Python ```python import requests import json api_key = "YOUR_API_KEY" url = "https://api.segmind.com/v1/happyhorse-1.1" data = { "prompt": "A majestic white horse galloping through golden wheat fields at sunset, cinematic slow motion, warm golden hour lighting, dust particles catching sunlight, shallow depth of field, professional nature documentary style", "image": "https://example.com/image.jpg", "negative_prompt": "blurry, low quality, distorted, watermark, text", "resolution": "1080p", "duration": 5, "aspect_ratio": "16:9", "seed": 0, "prompt_extend": true, "watermark": false, "reference_images": [] } response = requests.post( url, json=data, headers={ 'x-api-key': api_key, 'Content-Type': 'application/json' } ) if response.status_code == 200: # For image/video/audio models, response.content contains the binary data with open('output.png', 'wb') as f: f.write(response.content) print('Generation complete, saved to output.png') else: print(f"Error: {response.status_code}") print(response.text) ``` ### JavaScript ```javascript const apiKey = 'YOUR_API_KEY'; const url = 'https://api.segmind.com/v1/happyhorse-1.1'; const data = { "prompt": "A majestic white horse galloping through golden wheat fields at sunset, cinematic slow motion, warm golden hour lighting, dust particles catching sunlight, shallow depth of field, professional nature documentary style", "image": "https://example.com/image.jpg", "negative_prompt": "blurry, low quality, distorted, watermark, text", "resolution": "1080p", "duration": 5, "aspect_ratio": "16:9", "seed": 0, "prompt_extend": true, "watermark": false, "reference_images": [] }; const response = await fetch(url, { method: 'POST', headers: { 'x-api-key': apiKey, 'Content-Type': 'application/json', }, body: JSON.stringify(data), }); if (response.ok) { // For image/video/audio models, response contains binary data const blob = await response.blob(); const downloadUrl = URL.createObjectURL(blob); // Create download link const a = document.createElement('a'); a.href = downloadUrl; a.download = 'output.png'; a.click(); console.log('Generation complete'); } ``` ## Additional Resources ### Documentation - [Model Playground](https://www.segmind.com/models/happyhorse-1.1) - [API Documentation](https://www.segmind.com/models/happyhorse-1.1/api) - [Pricing Details](https://www.segmind.com/models/happyhorse-1.1/pricing) - [Platform Documentation](https://docs.segmind.com/)