Kokoro 82M — Open-Weight Text-to-Speech (TTS) API
What is Kokoro 82M?
Kokoro 82M is a lightweight, open-weight text-to-speech model that turns written text into natural, expressive speech. Despite having just 82 million parameters, it delivers voice quality comparable to models many times its size, while running dramatically faster and cheaper. Created by hexgrad and released under the permissive Apache 2.0 license, Kokoro is built on a decoder-only architecture combining StyleTTS 2 and ISTFTNet. On Segmind, you call it as a simple text-to-speech API: send text, pick a voice, and receive a ready-to-play WAV or MP3 audio file. It is the go-to choice for developers who need production-grade narration without the cost or footprint of large TTS systems.
Key Features
- •54 preset voices spanning American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese.
- •The voice ID also selects the language, so switching accents or languages is a single parameter change.
- •Adjustable speaking speed (0.5x–2.0x) for narration, dialogue, or fast notifications.
- •WAV (24kHz mono) or MP3 (128 kbps) output, with per-character billing on inputs up to 5,000 characters.
- •Extremely fast generation — up to roughly 96x real-time on GPU — ideal for high-volume batch voiceover.
Best Use Cases
Kokoro 82M excels at English narration for audiobooks, podcasts, YouTube voiceovers, and e-learning courses, where its American English voices sound clean and expressive. It is a strong fit for accessibility tools and screen readers, in-app voice prompts, IVR and notification systems, and voice agents that need low-latency speech. Because it is cost-efficient and fast, it is also ideal for prototyping and for high-throughput pipelines that convert large volumes of text to audio. In testing, a natural narrative sentence with the default af_heart voice produced clean, lifelike 24kHz speech with natural pauses and no clipping.
Prompt Tips and Output Quality
Write text the way you want it read: use punctuation to shape pacing, spell out ambiguous abbreviations, and split long scripts into separate calls, then stitch the audio. Choose a voice that matches your target language, since the voice determines the accent. American English voices are the most polished; non-English voices are usable but more variable. Keep speed near 1.0 for narration and raise it slightly for snappy UI prompts. Use WAV for the highest fidelity and MP3 when file size matters.

