194
models
Qwen Image 2 Models
Qwen Image 2 is Alibaba's second-generation multimodal image generation and editing suite — featuring a powerful lineup of text-to-image and instruction-following image editing models. The collection includes Qwen Image, Qwen Image 2512, Qwen Image Fast, Qwen Image Edit, and the comprehensive Qwen Image Edit Plus series with specialized tools: eraser, relighting, multi-LoRA, product photography, face-to-portrait, group photo generation, texture manipulation, next-scene generation, add-people, and blend-image capabilities. Qwen Image 2 models excel at understanding complex natural language editing instructions and executing them with high fidelity — preserving image structure while making targeted, intelligent modifications. The Edit Plus variants are especially powerful for creative and commercial workflows: apply multiple LoRA style adapters simultaneously, perform surgical object edits, generate product shots in new environments, and relight scenes — all from simple text prompts. On Segmind, every Qwen Image 2 model is available as a pay-per-use API endpoint. Chain editing models together in Segmind Workflows to build fully automated image production pipelines — generate, edit, relight, and upscale in one sequence.
Pruna P Video 2 Pro
Native-audio AI video from text or first/last-frame images.
Pruna P Video 2
Pruna's next-gen video model: text, image or audio in, 1080p video with native audio — dialogue, music and SFX — out.
GPT Image 2.5 Sunburst
Precisely edit and generate images with legible in-image text.
GPT Image 2.5 Flare
Fast text-to-image and editing with legible in-image text.
GPT 6 Astra
Complex reasoning, code generation, and image understanding; 1M-token context.
MiniMax H3 Max Turbo
Fast text-to-video and image-to-video, up to 15s at 768P.
Pruna P Video Edit
Edit video from a text prompt, keep original motion.
Gemini Omni 1.1
Text-to-video with synchronized native audio, up to 4K.
Lyria 3
Generate 30-second songs with vocals from text or images.
Wan 2.6 Image to Video Flash
Animate photos into 15-second 1080p video with native audio.
Grok Imagine Image 2
Text-to-image and image editing with crisp, legible text.
LTX 2.5 Fast
Text-to-video and image-to-video with native audio, up to 4K.
Qwen Image 3.0
Generate and edit legible in-image text, up to 2K.
Seedream 5.0 Pro Layer Decomposition
Split any image into editable transparent PNG layers.
Seedance 2.5
Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references.
FLUX 3 Image to Video
Animate images into 20-second clips with synchronized native audio.
Grok Imagine Video 1.5 Reference to Video
Character-consistent video from up to 7 reference images.
Grok Imagine Video 1.5 Image to Video
Animate a still image into 1080p video with synced audio.
MiniMax Hailuo H3 Image to Video
Animate a still image into 2K video up to 15s.
Pruna P Image Ideogram
Sub-second text-to-image with legible in-image text.
Ideogram V4 Remix
Restyle any image into posters with legible in-image text.
Ideogram V4 Fast
Generate posters and logos with accurate in-image text.
Seedream 5.0 Pro
Region-precise image editing with native multilingual text.
Higgsfield Soul 2.0
Generate fashion-editorial photorealistic photos from text or reference.
VEED Fabric 1.0
Animate any image into a realistic talking video, lip-synced to your audio or generated from a text script.
Pruna P Video Animate
Transfer video motion and audio onto any still image.
Nano Banana 2 Lite
Generate and edit 1K images in about four seconds.
Gemini Omni Flash
Text-to-video and image-to-video with synchronized native audio.
Pruna P Image Try-On
Dress photos in multiple garments with photorealistic virtual try-on.
Seedance 2.0 Mini
Fast text-to-video and image-to-video with synchronized audio.
HappyHorse 1.1
Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or reference images.
Luma Ray 3.2
Cinematic text-to-video and image-to-video clips up to 1080p.
Luma Uni-1 Max
Generate and edit images from plain-text instructions.
Luma Uni-1
Reasoning-first text-to-image and natural-language image editing.
Grok Imagine Video 1.5 (Preview)
Image-to-video with native synchronized audio, up to 720p.
Grok Imagine Video
Text-to-video and image-to-video with native synchronized audio.
Ideogram 4.0
Generate 2K posters and logos with accurate text rendering.
Grok Imagine Image
Text-to-image generation and editing, up to 2K resolution.
Pixverse Mimic
Transfer motion from reference videos onto still images.
Gemini Embedding 2
Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-specific modes.
Gemini 3.1 Pro
Frontier reasoning across text, images, video, and code.
Smart Banner Resizer
Recompose one image into multiple ad and banner sizes.
GPT Image 2
Generate photorealistic images with legible multilingual text and 2K output.
Claude Opus 4.7
Anthropic's most capable AI model excelling at agentic coding, complex reasoning, and high-resolution vision with a 1M-token context window.
Seedance 2.0 Fast
Professional-grade video creation model with native audio, similar to SeeDance 2.0 but faster and cheaper.
Seedance 2.0
Cinematic AI videos with native audio and multi-shot narratives.
Wan 2.7 Reference to Video
Character-consistent multi-subject videos from reference images.
Wan 2.7 Image to Video
Animate any image into cinematic 1080P video with audio.
Wan 2.7 Image Generation Pro
4K images with chain-of-thought reasoning and multilingual text.
Wan 2.7 Image Generation
2K image generation with precise multilingual text rendering.