164
models
Wan 2.1 Video Models
Wan 2.1 is Alibaba's landmark open-source video generation model series that set a new standard for accessible, production-ready AI video when it launched. Available in 480p and 720p resolution variants for image-to-video, plus a powerful text-to-video model, Wan 2.1 delivers smooth motion synthesis, accurate subject tracking, and natural scene dynamics across a wide range of content types. The models are built on large-scale training and an open architecture that became the foundation for subsequent Wan releases and community fine-tunes. Wan 2.1 is complemented by Wan Animate for targeted object and scene animation, Wan Scail for video upscaling workflows, and Wan Video Effects for creative visual transformations. The 480p variant is optimized for speed and throughput, while the 720p variant delivers higher visual fidelity for production use. Wan 2.1's open-weight design makes it particularly valuable for developers building custom video generation applications, researchers experimenting with video AI architectures, and production teams who need a reliable, well-tested video generation backbone. On Segmind, all Wan 2.1 models are available as pay-per-use APIs — no GPU management, no infrastructure overhead. Integrate professional video generation into your application with a single API call.
MiniMax Hailuo H3 Image to Video
Animate a still image into 2K video up to 15s.
Reve 2
Generate and edit 4K images with sharp in-image text.
VEED Lipsync v2
Dub talking-head videos with emotion-matched lip-sync.
MiniMax M3
Reason over 1M-token context for coding and agents.
Nemotron 3 Ultra
1M-token reasoning for coding agents and deep research.
GLM 5.2
1M-token open-weight LLM for long-horizon coding.
Higgsfield Soul 2.0
Generate fashion-editorial photorealistic photos from text or reference.
VEED Fabric 1.0
Animate any image into a realistic talking video, lip-synced to your audio or generated from a text script.
Nano Banana 2 Lite
Generate and edit 1K images in about four seconds.
Seed Audio 1.0
Generate full audio scenes: dialogue, music, effects, voice cloning.
Seedance 2.0 Mini
Fast text-to-video and image-to-video with synchronized audio.
HappyHorse 1.1
Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or reference images.
Luma Ray 3.2
Cinematic text-to-video and image-to-video clips up to 1080p.
Luma Uni-1 Max
Generate and edit images from plain-text instructions.
Luma Uni-1
Reasoning-first text-to-image and natural-language image editing.
Grok Imagine Video 1.5 (Preview)
Image-to-video with native synchronized audio, up to 720p.
Gemini 3.1 Flash TTS
Expressive, controllable TTS with 70+ language support.
Gemini Embedding 2
Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-specific modes.
Gemini Embedding 001
MTEB #1 text embeddings for RAG, search, and clustering.
Gemini 2.5 Flash Lite
Fastest Gemini 2.5 model for high-volume text and vision tasks.
Gemini 3.1 Flash Lite
Ultra-fast, affordable LLM for high-volume AI pipelines.
Gemini 3.1 Pro
Frontier reasoning across text, images, video, and code.
GPT 5.5
Frontier reasoning and coding with 1M-token context window.
HappyHorse 1.0
Cinematic 1080p text-to-video with native audio and lip-sync.
GPT Image 2
Generate photorealistic images with legible multilingual text and 2K output.
Claude Opus 4.7
Anthropic's most capable AI model excelling at agentic coding, complex reasoning, and high-resolution vision with a 1M-token context window.
Seedance 2.0 Fast
Professional-grade video creation model with native audio, similar to SeeDance 2.0 but faster and cheaper.
Seedance 2.0
Cinematic AI videos with native audio and multi-shot narratives.
Wan 2.7 Video Editing
Edit existing videos precisely using natural language text instructions.
Wan 2.7 Reference to Video
Character-consistent multi-subject videos from reference images.
Wan 2.7 Image to Video
Animate any image into cinematic 1080P video with audio.
Wan 2.7 Text to Video
1080P cinematic videos with audio sync and multi-shot control.
Wan 2.7 Image Generation Pro
4K images with chain-of-thought reasoning and multilingual text.
Wan 2.7 Image Generation
2K image generation with precise multilingual text rendering.
Pixverse V6
15-second AI videos with native audio and cinematic controls.
Qwen Flash
Fastest low-cost LLM with 1M context for high-volume tasks.
Qwen Plus
Mid-tier 1M context LLM for summarization and content tasks.
Qwen 3 Coder Flash
High-volume code generation with 1M token context window.
Qwen 3 Max
1T-parameter LLM with hybrid reasoning and 262K context.
Qwen 3.5 Plus
Multimodal 1M context AI for image, video, and text.
Kling V3 Image 2 Image
Transform images into photorealistic, production-ready visuals.
Kling O3 Text-to-Video
15-second cinematic AI videos with native audio.
Wan 2.2 Image to Video Flash
Convert a single image into a coherent dynamic video.
Nano Banana 2
Fast photorealistic images — ideal for marketing and ads.
Kling 3.0 Pro Image-to-Video
Animated 1080p videos from images with dynamic motion.
Kling 3.0 Standard Image-to-Video
Controlled cinematic 1080p videos from starting images.
Kling 3.0 Pro Text-to-Video
Cinematic 1080p videos with realistic audio from text.
Kling 3.0 Standard Text-to-Video
Stunning 1080p cinematic videos from simple text prompts.
Flux-2 Klein-4b
Sub-second photorealistic image generation and editing.
Flux-2 Klein-9b
Ultra-fast photorealistic image generation on consumer GPUs.