75
models
Segmind Models
Segmind's own catalogue of production-tuned AI models, built and hosted in-house rather than proxied from a third party. This is the largest single collection on the platform, spanning face swap and identity models (HyperSwap, FaceSwap V5, SD2.1 Faceswapper), virtual try-on and fashion models (SegFit, Video Try-On), image upscaling and restoration, background and object removal, text-to-speech and voice models (Veena TTS, Veena Max), and a broad set of Stable Diffusion pipelines with ControlNet, inpainting and LoRA support. Because these models run on Segmind's own inference stack, they are tuned for the two things that decide whether a model is usable in production: predictable latency and per-call cost. Many are serverless and scale to zero, so you pay only for the generations you actually make. The collection is the practical starting point for teams building image and video pipelines who need reliable building blocks rather than frontier headline models: a faceswap step, an upscale step, a background removal step, a TTS step. Every model is available as a pay-per-use API endpoint with no GPU provisioning, and all of them chain together inside Segmind Workflows to automate complete production pipelines end to end.
Kokoro 82M
Text-to-speech with 54 multilingual voices.
Smart Banner Resizer
Recompose one image into multiple ad and banner sizes.
Segmind Faceswap v5
Ultra-fast face and head swapping in images.
Wan Scail
Professional character animations from reference images.
Video Tryon V2
Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in high-quality, fully-preserved motion up to 50 seconds.
InfiniteTalk
Full-body animation from images synchronized perfectly to audio.
Video Tryon
Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in high-quality, fully-preserved motion up to 50 seconds.
Segmind SegFit v1.3
SegFit v1.3 enables hyper-realistic virtual try-ons, enhancing online fashion retail experiences without physical photoshoots.
Faceswap V3 Multifaceswap
Faceswap V3 Multifaceswap enables realistic face swapping in images, preserving lighting and expressions for professional results.
Segmind SegFit v1.2
SegFit v1.2 creates hyper-realistic virtual try-on images, transforming fashion retail engagement and conversion rates.
Segmind FaceSwap Comic v1
FaceSwap Comic v1 is an AI-powered face swapping model designed to blend real faces into illustrated or cartoon-style images while preserving the target’s artistic look. Ideal for personalized children’s storybooks and stylized content, it offers fine control over facial expression, realism, and stylistic adaptation.
Caricature Style
Transform everyday photos into lively, whimsical caricature illustrations that highlight individual features with playful exaggeration.
Segmind Relighting V2
Transform images with customizable, photorealistic lighting for unparalleled visual creativity and authenticity.
Ace Step Music
ACE-Step generates high-quality music rapidly, enhancing the creative process for developers and artists worldwide.
Chroma
Chroma is an open-source, 8.9B parameter text-to-image model (based on FLUX.1-schnell) designed for diverse and uncensored content generation, including anime, furry art, and photography.
Nomos Image Upscaler 4k
This upscaling model is ideal for enhancing amateur to professional photos, excelling with subjects like cats, hair, and party scenes. It handles both small (as low as 300px) and large images well, delivering sharp, clear results even when significantly resized.
Skin Contrast Upscaler
Enhances skin detail in images while preserving background quality for professional photography and art.
Supir Photo-Realistic Image Restoration
SUPIR restores and enhances images to stunning, photo-realistic quality with advanced AI techniques.
Dia (Text to Speech)
Dia by Nari Labs is an advanced open-weights TTS model that brings scripts to life with natural speech, emotions, and nonverbal cues. Easily control tone, voice, and delivery. Great alternative to ElevenLabs.
Segmind SceneCraft v0.1
SceneCraft transforms plain or existing product images into visually rich, photorealistic scenes. Whether starting from a white background or enhancing existing settings, it works seamlessly across furniture, home decor, and consumer goods.
Segmind SegFit v1.1
Segmind's Fashion and Immersive Try-on model. SegFIT offers effortless AI virtual try-on from just a product image. No models needed! Boost engagement & conversions with this flexible and fast try-on model.
Warmth of Jesus
Experience the viral "Warmth of Jesus" effect on PixVerse! Transform your images into heartwarming videos of Jesus embracing people.
Muscle Surge
Instantly add muscle and strength to your videos with Pixverse Muscle Surge effect!
Segmind Relighting
Prompts to auto-magically relight your images.
Segmind SegSwap v0.1
Swap Objects Instantly. The Segmind SegSwap v0.1 model enables dynamic and precise image editing by allowing users to remove, replace, or add objects and transfer patterns seamlessly within images.
Segmind Faceswap v4
Segmind FaceSwap v4 enables fast and precise face or head swapping between images with customizable options for style, output format, and image quality. Designed for creators and designers, it ensures natural-looking results with reproducibility through seed control for consistent outputs.
Wan Video Effects
Transform your videos with diverse video effects. Start creating captivating videos today.
AI Face Swap (image and video)
AI Face Swap: Effortlessly replace faces online. Fine-tune swaps with advanced controls for age, gender, and resolution.
Omini Control
OminiControl is an innovative framework that optimizes Diffusion Transformer models for versatile image generation tasks.
AI Product Photography
Elevate your product imagery with our AI-powered photography model. Create stunning, professional-quality photos that boost engagement and sales. Perfect for e-commerce and digital marketing.
Transparent Background Maker
Transform your images with Transparent Background Maker. Quickly remove backgrounds using AI technology, supporting PNG and JPG formats. Ideal for enhancing product photos and creating eye-catching graphics for social media and marketing.
Faceswap V3
Face Swap V3 is a cutting-edge tool that empowers you to seamlessly swap faces in images. With customizable features and advanced technology, you can achieve professional-quality results.
Face Detailer
Restore characters' faces to their original glory with Face Detailer. Enhance facial details, eliminate distortion, and upscale images for stunning results.
Video Stitch
Revolutionize your video editing with the Video Stitch Model. Seamlessly stitch clips, add captivating audio, and create professional-looking videos in minutes.
Simple Vector Flux Lora
Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions
Expression Editor
Expression Editor uses reference images to accurately generate new images with desired expressions. Perfect for digital art, memes, and marketing.
Esrgan Video Upscaler
ESRGAN Video Upscaler: Experience sharper, clearer 4k videos with ESRGAN. This AI-powered video upscaler boosts resolution and reduces artifacts, making your video content look its best. Best Topaz alternative.
Consistent Character With Pose
Create images of a given character in different poses
AI Product Photo Editor
AI Product Photo Editor leverages advanced image-based ML techniques to generate high-quality product visuals using text prompts, product images, and background images.
Video Faceswap
Video Faceswap is a powerful tool for creators, filmmakers, and meme enthusiasts. With this innovative technology, you can effortlessly replace faces in videos
Aura Flow
Largest completely open sourced flow-based generation model that is capable of text-to-image generation
Story Diffusion
Story Diffusion turns your written narratives into stunning image sequences.
Omni Zero
Omni-Zero: A diffusion pipeline for zero-shot stylized portrait creation.
LLAVA 1.6 7B
LLaVa translates images into text descriptions & captions.
Tooncrafter
Create videos from illustrated input images
V Express
V-Express lets you create portrait videos from single images.
SadTalker
Audio-based Lip Synchronization for Talking Head Video
Hallo
Hallo lets you create portrait videos from single images.
Relighting
Prompts to auto-magically relight your images.
Magic Eraser
LaMA Object Removal- AI Magic Eraser