150
models
Wan 2.2 Models
Wan 2.2 is Alibaba's previous-generation high-performance video model series, available in Fast and Flash variants for both text-to-video and image-to-video generation. Wan 2.2 Fast delivers optimized generation speed for production workflows requiring consistent throughput, while Wan 2.2 Flash provides ultra-fast generation ideal for rapid prototyping, creative iteration, and preview workflows before committing to higher-quality renders. Both variants maintain Wan's characteristic strength in motion quality, scene coherence, and prompt adherence — delivering professional-grade video output suitable for marketing content, social media, and creative production. Wan 2.2 also introduced the Flash tier that became a template for faster inference across the Wan model family. These models are a cost-effective choice for high-volume video generation workflows where speed and throughput are priorities, and continue to be widely used in production environments. On Segmind, all Wan 2.2 variants are available as pay-per-use APIs with no infrastructure overhead.
MiniMax H3 Max Multi-Angle
Orbit, crane and dolly a frozen photo in 3D.
Pruna P Video 2 Pro
Native-audio AI video from text or first/last-frame images.
Pruna P Video 2
Pruna's next-gen video model: text, image or audio in, 1080p video with native audio — dialogue, music and SFX — out.
GPT Image 2.5 Sunburst
Precisely edit and generate images with legible in-image text.
GPT Image 2.5 Flare
Fast text-to-image and editing with legible in-image text.
Seed Audio 2.0
Next-gen ByteDance audio model: video dubbing, translation, vocal/SFX/BGM stems, 6-minute output, 30 languages.
GPT 6 Astra
Complex reasoning, code generation, and image understanding; 1M-token context.
Wan 3.0 Video
Generate 30-second 1080p video with native audio.
Wan 2.6 Image to Video Flash
Animate photos into 15-second 1080p video with native audio.
Grok Imagine Image 2
Text-to-image and image editing with crisp, legible text.
LTX 2.5 Pro
Generate 1080p video with native audio and multi-shot scenes.
LTX 2.5 Fast
Text-to-video and image-to-video with native audio, up to 4K.
Qwen Image 3.0
Generate and edit legible in-image text, up to 2K.
Seedance 2.5
Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references.
FLUX 3 Image to Video
Animate images into 20-second clips with synchronized native audio.
FLUX 3 Text to Video
Cinematic text-to-video with native lip-synced audio, up to 20s.
Qwen3.8 Max
Multimodal reasoning and agentic coding with 1M-token context.
MiniMax Hailuo H3 Reference to Video
Keep characters and products consistent in 2K reference-to-video.
MiniMax Hailuo H3 Image to Video
Animate a still image into 2K video up to 15s.
MiniMax Hailuo H3 Text to Video
Text-to-video: cinematic 2K clips with native audio.
MiniMax M3
Reason over 1M-token context for coding and agents.
GLM 5.2
1M-token open-weight LLM for long-horizon coding.
Higgsfield Soul 2.0
Generate fashion-editorial photorealistic photos from text or reference.
VEED Avatars
Generate UGC-style talking avatar videos from text or audio using 28 stock presenters with realistic lip-sync.
Nano Banana 2 Lite
Generate and edit 1K images in about four seconds.
Pruna P Image Try-On
Dress photos in multiple garments with photorealistic virtual try-on.
Seedance 2.0 Mini
Fast text-to-video and image-to-video with synchronized audio.
Luma Ray 3.2
Cinematic text-to-video and image-to-video clips up to 1080p.
Grok Text-to-Speech
Convert text to speech in 20 languages with five voices.
Ideogram 4.0
Generate 2K posters and logos with accurate text rendering.
Grok Imagine Image
Text-to-image generation and editing, up to 2K resolution.
Gemini Embedding 2
Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-specific modes.
Gemini 2.5 Flash Lite
Fastest Gemini 2.5 model for high-volume text and vision tasks.
Smart Banner Resizer
Recompose one image into multiple ad and banner sizes.
GPT Image 2
Generate photorealistic images with legible multilingual text and 2K output.
Claude Opus 4.7
Anthropic's most capable AI model excelling at agentic coding, complex reasoning, and high-resolution vision with a 1M-token context window.
Seedance 2.0 Fast
Professional-grade video creation model with native audio, similar to SeeDance 2.0 but faster and cheaper.
Seedance 2.0
Cinematic AI videos with native audio and multi-shot narratives.
Wan 2.7 Video Editing
Edit existing videos precisely using natural language text instructions.
Wan 2.7 Reference to Video
Character-consistent multi-subject videos from reference images.
Wan 2.7 Image to Video
Animate any image into cinematic 1080P video with audio.
Wan 2.7 Text to Video
1080P cinematic videos with audio sync and multi-shot control.
Wan 2.7 Image Generation Pro
4K images with chain-of-thought reasoning and multilingual text.
Wan 2.7 Image Generation
2K image generation with precise multilingual text rendering.
Qwen 3 VL Flash
Fast, affordable vision-language model with 262K context OCR.
Qwen 3 Max
1T-parameter LLM with hybrid reasoning and 262K context.
Kling V3 Image 2 Image
Transform images into photorealistic, production-ready visuals.
Kling V3 Text to Image
Photorealistic, print-ready images from text prompts.
Wan 2.2 Image to Video Flash
Convert a single image into a coherent dynamic video.
Nano Banana 2
Fast photorealistic images — ideal for marketing and ads.