194
models
Qwen Image 2 Models
Qwen Image 2 is Alibaba's second-generation multimodal image generation and editing suite — featuring a powerful lineup of text-to-image and instruction-following image editing models. The collection includes Qwen Image, Qwen Image 2512, Qwen Image Fast, Qwen Image Edit, and the comprehensive Qwen Image Edit Plus series with specialized tools: eraser, relighting, multi-LoRA, product photography, face-to-portrait, group photo generation, texture manipulation, next-scene generation, add-people, and blend-image capabilities. Qwen Image 2 models excel at understanding complex natural language editing instructions and executing them with high fidelity — preserving image structure while making targeted, intelligent modifications. The Edit Plus variants are especially powerful for creative and commercial workflows: apply multiple LoRA style adapters simultaneously, perform surgical object edits, generate product shots in new environments, and relight scenes — all from simple text prompts. On Segmind, every Qwen Image 2 model is available as a pay-per-use API endpoint. Chain editing models together in Segmind Workflows to build fully automated image production pipelines — generate, edit, relight, and upscale in one sequence.
MiniMax Hailuo H3 Image to Video
Animate a still image into 2K video up to 15s.
Pruna P Image Ideogram
Sub-second text-to-image with legible in-image text.
Reve 2
Generate and edit 4K images with sharp in-image text.
Ideogram V4 Remix
Restyle any image into posters with legible in-image text.
Ideogram V4 Fast
Generate posters and logos with accurate in-image text.
Seedream 5.0 Pro
Region-precise image editing with native multilingual text.
VEED Fabric 1.0
Animate any image into a realistic talking video, lip-synced to your audio or generated from a text script.
Pruna P Video Animate
Transfer video motion and audio onto any still image.
Nano Banana 2 Lite
Generate and edit 1K images in about four seconds.
Gemini Omni Flash
Text-to-video and image-to-video with synchronized native audio.
Pruna P Image Try-On
Dress photos in multiple garments with photorealistic virtual try-on.
Seedance 2.0 Mini
Fast text-to-video and image-to-video with synchronized audio.
HappyHorse 1.1
Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or reference images.
Luma Ray 3.2
Cinematic text-to-video and image-to-video clips up to 1080p.
Luma Uni-1 Max
Generate and edit images from plain-text instructions.
Luma Uni-1
Reasoning-first text-to-image and natural-language image editing.
Grok Imagine Video 1.5 (Preview)
Image-to-video with native synchronized audio, up to 720p.
Grok Imagine Video
Text-to-video and image-to-video with native synchronized audio.
Grok Imagine Image
Text-to-image generation and editing, up to 2K resolution.
Pixverse Mimic
Transfer motion from reference videos onto still images.
Gemini Embedding 2
Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-specific modes.
Imagen 4 Fast
Fast photorealistic image generation for bulk and iteration.
Imagen 4 Ultra
Photorealistic images with native 2K resolution and precise text.
Gemini 3.1 Pro
Frontier reasoning across text, images, video, and code.
Smart Banner Resizer
Recompose one image into multiple ad and banner sizes.
GPT Image 2
Generate photorealistic images with legible multilingual text and 2K output.
Wan 2.7 Reference to Video
Character-consistent multi-subject videos from reference images.
Wan 2.7 Image to Video
Animate any image into cinematic 1080P video with audio.
Wan 2.7 Image Generation Pro
4K images with chain-of-thought reasoning and multilingual text.
Wan 2.7 Image Generation
2K image generation with precise multilingual text rendering.
Qwen Flash
Fastest low-cost LLM with 1M context for high-volume tasks.
Qwen Plus
Mid-tier 1M context LLM for summarization and content tasks.
Qwen 3 VL Flash
Fast, affordable vision-language model with 262K context OCR.
Qwen 3 VL Plus
Powerful visual QA and document analysis from images.
Qwen 3 Coder Flash
High-volume code generation with 1M token context window.
Qwen 3 Coder Plus
Generates, debugs, and refactors entire codebases efficiently.
Qwen 3 Max
1T-parameter LLM with hybrid reasoning and 262K context.
Qwen 3.5 Plus
Multimodal 1M context AI for image, video, and text.
Qwen 3.5 Flash
Fast multimodal AI processing text, images, and video affordably.
HyperSwap Image Faceswap by FaceFusion Labs
High-quality face swapping built for real production workflows.
Kling O3 Image To Video
Images to cinematic videos with precise motion control.
Kling V3 Image 2 Image
Transform images into photorealistic, production-ready visuals.
Kling V3 Text to Image
Photorealistic, print-ready images from text prompts.
Kling O3 Video To Video Reference
Swap characters and restyle videos using reference images.
HyperSwap: Video Faceswap by FaceFusion Labs
Realistic face swapping in videos from a single image.
Wan 2.2 Image to Video Flash
Convert a single image into a coherent dynamic video.
Seedream 5.0 Lite: Image-to-Image
Transform images intelligently with detailed text prompts.
Seedream 5.0 Lite: Text-to-Image
Fast, affordable instruction-following image generation.
Nano Banana 2
Fast photorealistic images — ideal for marketing and ads.
Kling 3.0 Pro Image-to-Video
Animated 1080p videos from images with dynamic motion.