LTX 2.5 Fast — Image-to-Video and Text-to-Video Generation
What is LTX 2.5 Fast?
LTX 2.5 Fast is the speed-optimized variant of LTX-2.5, the open-weights video foundation model from LTX (the world-model company spun out of Lightricks). It turns a text prompt or a still image into a finished clip with synchronized native audio, in portrait or landscape, at resolutions up to 4K and lengths up to 20 seconds. Where the Pro variant targets maximum fidelity at 1080p, Fast is built for rapid iteration and production-volume generation — the same generation quality with fewer retries and far lower latency, so you can batch variations and refresh creative quickly.
Because both text-to-video and image-to-video run through one endpoint, you can start from a written scene or animate an existing frame. Add a closing frame with last_frame_uri to control exactly where an image-to-video clip lands.
Key Features
- •Native audio: generate a synchronized soundtrack — ambience, music, dialogue, or singing — described directly in your prompt.
- •Multi-shot scenes: a single generation renders several connected shots that hold character, scene, lighting, style, and voice across cuts.
- •Up to 4K, up to 20s: 720p, 1080p, 2K and 4K output in 16:9 or 9:16.
- •Automatic duration: let the model choose the clip length from the described action.
- •Camera control: force a move (dolly, jib, focus shift) or leave it on auto to let the prompt drive the shot.
- •Cinematic frame rates: 24/25 fps for a filmic look, 48/50 fps for smoother motion.
Best Use Cases
LTX 2.5 Fast fits short-form social video for TikTok, Reels and Shorts, product and ad creative, previsualization, and explainer content. Its multi-shot consistency makes it usable for campaign sequences that need a character to hold shot to shot, and its speed makes overnight batch generation and rapid A/B iteration practical rather than theoretical.
Prompt Tips and Output Quality
Write present-tense, single-subject scenes of roughly 4–8 sentences. Establish the shot, set lighting and atmosphere, describe the action as one flowing sequence, and describe the audio — place spoken dialogue in quotation marks. For multi-shot prompts, name each cut (hard cut, match cut, dissolve), re-establish the new framing, and state whether music or dialogue carries across. Keep on-screen text short and prominent; exact spelling across frames is not guaranteed, so add critical titles or logos in post.
FAQs
Is LTX 2.5 Fast text-to-video or image-to-video? Both. Provide a prompt alone for text-to-video, or add an image to switch to image-to-video.
Does it generate audio? Yes — a synchronized audio track is produced by default. Set generate_audio to false for silent video.
What is the maximum resolution and length? Up to 4K and up to 20 seconds, though 2K, 4K and 48/50 fps cap duration at 10 seconds.
How do multi-shot scenes work? One prompt can describe several shots joined by explicit cuts; the model keeps characters and style consistent across them.
How is Fast different from Pro? Fast reaches up to 4K and prioritizes speed and low latency; Pro tops out at 1080p and prioritizes fidelity.
Can I control where a clip ends? Yes, on image-to-video supply last_frame_uri — but it cannot be combined with automatic duration.