81
models
Segmind Models
Segmind's own catalogue of production-tuned AI models, built and hosted in-house rather than proxied from a third party. This is the largest single collection on the platform, spanning face swap and identity models (HyperSwap, FaceSwap V5, SD2.1 Faceswapper), virtual try-on and fashion models (SegFit, Video Try-On), image upscaling and restoration, background and object removal, text-to-speech and voice models (Veena TTS, Veena Max), and a broad set of Stable Diffusion pipelines with ControlNet, inpainting and LoRA support. Because these models run on Segmind's own inference stack, they are tuned for the two things that decide whether a model is usable in production: predictable latency and per-call cost. Many are serverless and scale to zero, so you pay only for the generations you actually make. The collection is the practical starting point for teams building image and video pipelines who need reliable building blocks rather than frontier headline models: a faceswap step, an upscale step, a background removal step, a TTS step. Every model is available as a pay-per-use API endpoint with no GPU provisioning, and all of them chain together inside Segmind Workflows to automate complete production pipelines end to end.
Kokoro 82M
Text-to-speech with 54 multilingual voices.
Smart Banner Resizer
Recompose one image into multiple ad and banner sizes.
Segmind Faceswap v5
Ultra-fast face and head swapping in images.
Wan Scail
Professional character animations from reference images.
Z Image Turbo
Photorealistic images in under one second, bilingual text.
Video Tryon V2
Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in high-quality, fully-preserved motion up to 50 seconds.
InfiniteTalk
Full-body animation from images synchronized perfectly to audio.
Wan Animate
Animate characters and replace video subjects seamlessly.
Video Tryon
Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in high-quality, fully-preserved motion up to 50 seconds.
Segmind SegFit v1.3
SegFit v1.3 enables hyper-realistic virtual try-ons, enhancing online fashion retail experiences without physical photoshoots.
Infinite You
InfiniteYou generates high-fidelity portraits preserving identity while aligning with creative text prompts.
Faceswap V3 Multifaceswap
Faceswap V3 Multifaceswap enables realistic face swapping in images, preserving lighting and expressions for professional results.
Veena TTS
Veena transforms text into high-fidelity, expressive speech in Hindi and English for real-time applications.
Segmind SegFit v1.2
SegFit v1.2 creates hyper-realistic virtual try-on images, transforming fashion retail engagement and conversion rates.
Segmind FaceSwap Comic v1
FaceSwap Comic v1 is an AI-powered face swapping model designed to blend real faces into illustrated or cartoon-style images while preserving the target’s artistic look. Ideal for personalized children’s storybooks and stylized content, it offers fine control over facial expression, realism, and stylistic adaptation.
Caricature Style
Transform everyday photos into lively, whimsical caricature illustrations that highlight individual features with playful exaggeration.
Segmind Relighting V2
Transform images with customizable, photorealistic lighting for unparalleled visual creativity and authenticity.
Ace Step Music
ACE-Step generates high-quality music rapidly, enhancing the creative process for developers and artists worldwide.
Chroma
Chroma is an open-source, 8.9B parameter text-to-image model (based on FLUX.1-schnell) designed for diverse and uncensored content generation, including anime, furry art, and photography.
Nomos Image Upscaler 4k
This upscaling model is ideal for enhancing amateur to professional photos, excelling with subjects like cats, hair, and party scenes. It handles both small (as low as 300px) and large images well, delivering sharp, clear results even when significantly resized.
Skin Contrast Upscaler
Enhances skin detail in images while preserving background quality for professional photography and art.
Supir Photo-Realistic Image Restoration
SUPIR restores and enhances images to stunning, photo-realistic quality with advanced AI techniques.
Dia (Text to Speech)
Dia by Nari Labs is an advanced open-weights TTS model that brings scripts to life with natural speech, emotions, and nonverbal cues. Easily control tone, voice, and delivery. Great alternative to ElevenLabs.
Segmind SceneCraft v0.1
SceneCraft transforms plain or existing product images into visually rich, photorealistic scenes. Whether starting from a white background or enhancing existing settings, it works seamlessly across furniture, home decor, and consumer goods.
Segmind SegFit v1.1
Segmind's Fashion and Immersive Try-on model. SegFIT offers effortless AI virtual try-on from just a product image. No models needed! Boost engagement & conversions with this flexible and fast try-on model.
Warmth of Jesus
Experience the viral "Warmth of Jesus" effect on PixVerse! Transform your images into heartwarming videos of Jesus embracing people.
Muscle Surge
Instantly add muscle and strength to your videos with Pixverse Muscle Surge effect!
Segmind Relighting
Prompts to auto-magically relight your images.
Segmind SegSwap v0.1
Swap Objects Instantly. The Segmind SegSwap v0.1 model enables dynamic and precise image editing by allowing users to remove, replace, or add objects and transfer patterns seamlessly within images.
Segmind Faceswap v4
Segmind FaceSwap v4 enables fast and precise face or head swapping between images with customizable options for style, output format, and image quality. Designed for creators and designers, it ensures natural-looking results with reproducibility through seed control for consistent outputs.
Wan Video Effects
Transform your videos with diverse video effects. Start creating captivating videos today.
AI Face Swap (image and video)
AI Face Swap: Effortlessly replace faces online. Fine-tune swaps with advanced controls for age, gender, and resolution.
Omini Control
OminiControl is an innovative framework that optimizes Diffusion Transformer models for versatile image generation tasks.
AI Product Photography
Elevate your product imagery with our AI-powered photography model. Create stunning, professional-quality photos that boost engagement and sales. Perfect for e-commerce and digital marketing.
Transparent Background Maker
Transform your images with Transparent Background Maker. Quickly remove backgrounds using AI technology, supporting PNG and JPG formats. Ideal for enhancing product photos and creating eye-catching graphics for social media and marketing.
Faceswap V3
Face Swap V3 is a cutting-edge tool that empowers you to seamlessly swap faces in images. With customizable features and advanced technology, you can achieve professional-quality results.
Face Detailer
Restore characters' faces to their original glory with Face Detailer. Enhance facial details, eliminate distortion, and upscale images for stunning results.
Video Stitch
Revolutionize your video editing with the Video Stitch Model. Seamlessly stitch clips, add captivating audio, and create professional-looking videos in minutes.
Simple Vector Flux Lora
Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Expression Editor
Expression Editor uses reference images to accurately generate new images with desired expressions. Perfect for digital art, memes, and marketing.
Esrgan Video Upscaler
ESRGAN Video Upscaler: Experience sharper, clearer 4k videos with ESRGAN. This AI-powered video upscaler boosts resolution and reduces artifacts, making your video content look its best. Best Topaz alternative.
Consistent Character With Pose
Create images of a given character in different poses
Fast Flux.1 Schnell
Fast Flux.1 Schnell by Segmind is an optimized text-to-image model designed for developers needing faster image generation. It offers high efficiency without compromising quality. Perfect for startups and engineers seeking quick, resource-efficient AI models.
AI Product Photo Editor
AI Product Photo Editor leverages advanced image-based ML techniques to generate high-quality product visuals using text prompts, product images, and background images.
Live Portrait video to video
Experience the magic of Live Portrait’s Video-to-Video Model! Transform your static images into dynamic videos seamlessly.
Video Faceswap
Video Faceswap is a powerful tool for creators, filmmakers, and meme enthusiasts. With this innovative technology, you can effortlessly replace faces in videos
Aura Flow
Largest completely open sourced flow-based generation model that is capable of text-to-image generation
Live Portrait
Live Portrait animates static images using a reference driving video through implicit key point based framework, bringing a portrait to life with realistic expressions and movements. It identifies key points on the face (think eyes, nose, mouth) and manipulates them to create expressions and movements.
Kolors
Kolors is a cutting-edge text-to-image model that bridges language and visual art. Transform your textual ideas into photorealistic images with semantic precision.
Story Diffusion
Story Diffusion turns your written narratives into stunning image sequences.