Gemini 3.8 Flash TTS — Text-to-Speech API
What is Gemini 3.8 Flash TTS?
Gemini 3.8 Flash TTS is Google's flagship creative text-to-speech model, built to perform a script rather than just read it. Instead of picking a fixed preset, you supply a transcript, choose a studio voice, and direct the delivery in plain English — pacing, emotion, accent, and dialect all shift line by line. It is designed for developers and creators building audiobooks, podcasts, games, e-learning, and voice agents who need speech that sounds acted, not announced.
The model treats your text as a verbatim transcript and keeps all delivery control in a separate style field, so your prompt stays clean and predictable. It returns studio-quality WAV audio and supports automatic language detection across 130+ languages.
Key Features
- •Line-by-line style direction: steer tone, pacing, emotion, and accent with a natural-language
styleprompt. - •Inline vocal cues: drop
<laughs>,<sigh>,<gasp>, or<short pause>exactly where they belong in the transcript. - •Native two-speaker dialogue: stage a two-person scene from a single call with distinct voices and natural turn-taking.
- •30 studio voices: expressive named voices such as Kore, Puck, Fenrir, Aoede, and Enceladus.
- •Multilingual: 130+ languages with authentic regional accents, auto-detected from the transcript.
- •Expressiveness control: a
temperaturedial from steady narration to animated performance.
Best Use Cases
Reach for Gemini 3.8 Flash TTS when delivery matters: character voices and emotional beats for games and immersive audiobooks, two-host podcasts and dramatic scenes, warm multilingual narration for explainers and localization, and natural-sounding voice agents. In testing, a single-voice sports call swung from a hushed build-up to a euphoric eruption, a two-speaker lab scene traded four turns in distinct voices, a Mexican-Spanish market read nailed the accent from the transcript alone, and the same engine dropped to an intimate whisper purely through the style field.
Prompt Tips and Output Quality
Keep stage directions out of text — it is spoken verbatim — and use the style field for emotion and the inline <angle-bracket> cues for sounds. Set voice_2 to a different voice for dialogue and prefix each line with the speaker name. Raise temperature to roughly 1.0–1.4 for expressive scenes; keep it near 0.3 for steady reads. Outputs are clean 24 kHz mono WAV.