186
models
Vidu Models
Vidu is Shengshu Technology's advanced AI video generation model, distinguished by its exceptional subject consistency and creative control. The Vidu collection includes Vidu Q1 Reference to Video — which generates high-quality video from a reference image while maintaining precise subject identity and appearance throughout the sequence, making it ideal for character-driven content, product showcases, and any workflow where visual consistency matters — and Vidu Template for standardized, template-driven video creation that enables repeatable, on-brand content at scale. Vidu models are particularly strong at human motion quality, facial expression rendering, and natural body dynamics — producing videos where subjects look and move authentically rather than like AI-generated approximations. The reference-to-video capability is especially valuable for e-commerce brands (animate product photos consistently), content creators (maintain character identity across multiple videos), and marketing teams (produce personalized video content from a single reference image at scale). On Segmind, all Vidu models are available as pay-per-use APIs — generate reference-consistent video with a single endpoint call, no subscriptions or model hosting required. Integrate into Segmind Workflows to automate complete video production pipelines from image to finished video.
FLUX 3 Video Edit
Edit any clip with one prompt, motion kept intact.
MiniMax Hailuo H3 Max Reference to Video
Keep characters and products consistent across reference-to-video clips.
Pruna P Video 2 Pro
Native-audio AI video from text or first/last-frame images.
Pruna P Video 2
Pruna's next-gen video model: text, image or audio in, 1080p video with native audio — dialogue, music and SFX — out.
Seed Audio 2.0
Next-gen ByteDance audio model: video dubbing, translation, vocal/SFX/BGM stems, 6-minute output, 30 languages.
MiniMax H3 Max Turbo
Fast text-to-video and image-to-video, up to 15s at 768P.
Pruna P Video Edit
Edit video from a text prompt, keep original motion.
Gemini Omni 1.1
Text-to-video with synchronized native audio, up to 4K.
Gemini Omni 1.1 Video Extend
Extend short video clips into longer seamless scenes.
Gemini Omni 1.1 Video Edit
Edit videos with a text prompt, subject preserved.
Wan 3.0 Video
Generate 30-second 1080p video with native audio.
Wan 2.6 Image to Video Flash
Animate photos into 15-second 1080p video with native audio.
LTX 2.5 Pro
Generate 1080p video with native audio and multi-shot scenes.
LTX 2.5 Fast
Text-to-video and image-to-video with native audio, up to 4K.
Seedance 2.5
Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references.
FLUX 3 Draft Enhance
Upscale AI video drafts to Full-HD with native audio.
FLUX 3 Extend Video
Extend clips into seamless video continuations with synchronized audio.
FLUX 3 Image to Video
Animate images into 20-second clips with synchronized native audio.
FLUX 3 Text to Video
Cinematic text-to-video with native lip-synced audio, up to 20s.
Grok Imagine Video 1.5 Reference to Video
Character-consistent video from up to 7 reference images.
Grok Imagine Video 1.5 Image to Video
Animate a still image into 1080p video with synced audio.
Grok Imagine Video 1.5 Text to Video
Text-to-video clips up to 1080p with native synchronized audio.
Sonilo Video to Video
Add frame-synced AI music and sound effects to video.
Sonilo Video to Audio
Generate video-synced music and sound effects from footage.
MiniMax Hailuo H3 Reference to Video
Keep characters and products consistent in 2K reference-to-video.
MiniMax Hailuo H3 Image to Video
Animate a still image into 2K video up to 15s.
MiniMax Hailuo H3 Text to Video
Text-to-video: cinematic 2K clips with native audio.
VEED Lipsync v2
Dub talking-head videos with emotion-matched lip-sync.
VEED Subtitles
Automatically transcribes and burns styled, translated subtitles into any video with 30 presets and a single API call.
VEED Video Background Removal
Remove any video's background with no green screen, or cleanly key chroma footage, using AI matting.
VEED Avatars
Generate UGC-style talking avatar videos from text or audio using 28 stock presenters with realistic lip-sync.
VEED Lipsync
Re-syncs the lips of any talking-head video to a new speech audio track for realistic dubbing and localization.
VEED Fabric 1.0
Animate any image into a realistic talking video, lip-synced to your audio or generated from a text script.
OpusClip - Clips From Video
Turn long videos into captioned vertical shorts.
Pruna P Video Replace
Swap on-screen video characters while preserving motion and audio.
Pruna P Video Animate
Transfer video motion and audio onto any still image.
Gemini Omni Flash
Text-to-video and image-to-video with synchronized native audio.
Pruna P Video Avatar
Animate any portrait into a lip-synced talking avatar.
Seedance 2.0 Mini
Fast text-to-video and image-to-video with synchronized audio.
HappyHorse 1.1
Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or reference images.
Luma Ray 3.2
Cinematic text-to-video and image-to-video clips up to 1080p.
Grok Imagine Video 1.5 (Preview)
Image-to-video with native synchronized audio, up to 720p.
Grok Imagine Video
Text-to-video and image-to-video with native synchronized audio.
HeyGen Avatar V — Create Avatar
Train a Digital Twin avatar from reference video.
HeyGen Avatar V
Studio-quality talking-avatar videos from text or audio.
Pixverse Mimic
Transfer motion from reference videos onto still images.
Gemini Embedding 2
Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-specific modes.
Gemini 3.1 Pro
Frontier reasoning across text, images, video, and code.
HappyHorse 1.0
Cinematic 1080p text-to-video with native audio and lip-sync.
Seedance 2.0 Fast
Professional-grade video creation model with native audio, similar to SeeDance 2.0 but faster and cheaper.