197
models
Wan 2.1 Video Models
Wan 2.1 is Alibaba's landmark open-source video generation model series that set a new standard for accessible, production-ready AI video when it launched. Available in 480p and 720p resolution variants for image-to-video, plus a powerful text-to-video model, Wan 2.1 delivers smooth motion synthesis, accurate subject tracking, and natural scene dynamics across a wide range of content types. The models are built on large-scale training and an open architecture that became the foundation for subsequent Wan releases and community fine-tunes. Wan 2.1 is complemented by Wan Animate for targeted object and scene animation, Wan Scail for video upscaling workflows, and Wan Video Effects for creative visual transformations. The 480p variant is optimized for speed and throughput, while the 720p variant delivers higher visual fidelity for production use. Wan 2.1's open-weight design makes it particularly valuable for developers building custom video generation applications, researchers experimenting with video AI architectures, and production teams who need a reliable, well-tested video generation backbone. On Segmind, all Wan 2.1 models are available as pay-per-use APIs — no GPU management, no infrastructure overhead. Integrate professional video generation into your application with a single API call.
FLUX 3 Video Edit
Edit any clip with one prompt, motion kept intact.
MiniMax Hailuo H3 Max Reference to Video
Keep characters and products consistent across reference-to-video clips.
MiniMax H3 Max Multi-Angle
Orbit, crane and dolly a frozen photo in 3D.
Pruna P Video 2 Pro
Native-audio AI video from text or first/last-frame images.
Claude Fable 5.1
Coding, debugging, and long-horizon reasoning for agents.
Claude Sonnet 5
Agentic coding, debugging and reasoning with 1M-token context.
Claude Opus 5
Frontier coding, debugging, and agents with 1M-token context.
Pruna P Video 2
Pruna's next-gen video model: text, image or audio in, 1080p video with native audio — dialogue, music and SFX — out.
GPT Image 2.5 Sunburst
Precisely edit and generate images with legible in-image text.
GPT Image 2.5 Flare
Fast text-to-image and editing with legible in-image text.
Seed Audio 2.0
Next-gen ByteDance audio model: video dubbing, translation, vocal/SFX/BGM stems, 6-minute output, 30 languages.
GPT 6 Astra
Complex reasoning, code generation, and image understanding; 1M-token context.
MiniMax H3 Max Turbo
Fast text-to-video and image-to-video, up to 15s at 768P.
Sarvam Bulbul v3 TTS
Text-to-speech in 11 Indian languages with 37 voices.
Pruna P Video Edit
Edit video from a text prompt, keep original motion.
Gemini Omni 1.1
Text-to-video with synchronized native audio, up to 4K.
Gemini Omni 1.1 Video Extend
Extend short video clips into longer seamless scenes.
Gemini Omni 1.1 Video Edit
Edit videos with a text prompt, subject preserved.
Lyria 3 Pro
Full-length text-to-music songs with vocals and lyrics.
Lyria 3
Generate 30-second songs with vocals from text or images.
Gemini 3.7 Flash
Fast multimodal LLM for coding, agents, and long-document analysis.
Wan 3.0 Video
Generate 30-second 1080p video with native audio.
Wan 2.6 Image to Video Flash
Animate photos into 15-second 1080p video with native audio.
Grok Imagine Image 2
Text-to-image and image editing with crisp, legible text.
LTX 2.5 Pro
Generate 1080p video with native audio and multi-shot scenes.
LTX 2.5 Fast
Text-to-video and image-to-video with native audio, up to 4K.
Qwen Image 3.0
Generate and edit legible in-image text, up to 2K.
Seedream 5.0 Pro Layer Decomposition
Split any image into editable transparent PNG layers.
Bria Extract Object
Extract any named object into a transparent PNG cutout.
Seedance 2.5
Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references.
FLUX 3 Extend Video
Extend clips into seamless video continuations with synchronized audio.
FLUX 3 Image to Video
Animate images into 20-second clips with synchronized native audio.
FLUX 3 Text to Video
Cinematic text-to-video with native lip-synced audio, up to 20s.
Qwen3.8 Max
Multimodal reasoning and agentic coding with 1M-token context.
Grok Imagine Video 1.5 Reference to Video
Character-consistent video from up to 7 reference images.
Grok Imagine Video 1.5 Image to Video
Animate a still image into 1080p video with synced audio.
Grok Imagine Video 1.5 Text to Video
Text-to-video clips up to 1080p with native synchronized audio.
Sonilo Text to Audio
Commercial-safe music and sound effects from text prompts.
MiniMax Hailuo H3 Reference to Video
Keep characters and products consistent in 2K reference-to-video.
MiniMax Hailuo H3 Image to Video
Animate a still image into 2K video up to 15s.
MiniMax Hailuo H3 Text to Video
Text-to-video: cinematic 2K clips with native audio.
Pruna P Image Ideogram
Sub-second text-to-image with legible in-image text.
Ideogram V4 Remix
Restyle any image into posters with legible in-image text.
ElevenLabs Music
Generate full songs with vocals or instrumental from text.
MiniMax M3
Reason over 1M-token context for coding and agents.
Nemotron 3 Ultra
1M-token reasoning for coding agents and deep research.
GLM 5.2
1M-token open-weight LLM for long-horizon coding.
Ideogram V4 Fast
Generate posters and logos with accurate in-image text.
Seedream 5.0 Pro
Region-precise image editing with native multilingual text.
Higgsfield Soul 2.0
Generate fashion-editorial photorealistic photos from text or reference.