2

models

Resemble AI Voice Models

Resemble AI builds voice cloning and speech synthesis models that reproduce a specific person's voice from a short reference recording, along with the detection tooling to identify synthetic speech. The collection covers voice cloning, which captures timbre, accent, cadence and delivery from a sample and then speaks arbitrary new text in that voice; general text-to-speech for high-quality synthetic narration; and audio deepfake detection for verifying whether a given clip was machine-generated. The cloning models are what make personalised and localised audio practical at scale: one recording session yields a voice that can narrate an unlimited catalogue, be updated without re-recording, and speak languages the original speaker does not. That serves audiobook and course production, per-user personalised audio, brand voices that stay consistent across years of content, and accessibility work. The detection side is the necessary counterpart: as synthetic voice becomes routine, platforms and publishers need a way to verify what they are receiving. On Segmind, Resemble models are pay-per-use API endpoints. Chain them with avatar, lipsync and video models in Segmind Workflows to produce complete narrated video from a script in one automated run.