Gemini 3.8 Flash-Lite TTS

Expressive multilingual text-to-speech with two-speaker dialogue.

Playground
APIPricing
~8.75s
Example output
0:00 / 0:00

Gemini 3.8 Flash-Lite TTS — AI Text-to-Speech (Audio Generation)

What is Gemini 3.8 Flash-Lite TTS?

Gemini 3.8 Flash-Lite TTS is Google's high-throughput text-to-speech model in the Gemini Audio family. It turns written scripts into expressive, natural-sounding speech at scale, and instead of reading a script flatly it performs it: you direct delivery in plain English and shift tone, pacing, and emotion line by line. It offers 30 named voices, native two-speaker dialogue, inline vocal cues, and more than 100 languages. On Hume AI's Overall Quality Index it ranks second, behind only the flagship Gemini 3.8 Flash TTS, with major long-form stability gains over Gemini 3.1 Flash TTS.

Key Features

  • •Directable expressiveness — a plain-English style field and a temperature control shift register, pacing, and emotion across the whole read.
  • •Inline vocal cues — drop <laughs>, <sigh>, <gasp>, and <short pause> into the text and they render where you place them.
  • •Native two-speaker dialogue — set Voice 1 and Voice 2 and prefix each line to stage a two-person conversation from one script.
  • •30 voices, 100+ languages — strong blind-preference results in Brazilian Portuguese, Japanese, Hindi, Mexican Spanish, Vietnamese, and Modern Standard Arabic.
  • •Long-form stability — consistent timbre and natural pacing across extended passages with minimal speaker drift.

Best Use Cases

Gemini 3.8 Flash-Lite TTS fits high-volume dubbing and media localization, audiobook and podcast narration, read-aloud features, and expressive real-time voice agents. Two-speaker staging makes interviews, explainers, and dramatic scenes simple to produce from a single script.

Prompt Tips and Output Quality

Write natural sentences and use the style field for direction such as "cheerful and friendly" or "whispered urgently". Keep numbers and proper nouns in context; the model pronounces them cleanly. Raise temperature to 0.8–1.4 for emotional delivery and lower it to 0.2–0.5 for steady narration. For other languages, write the script natively in that language. Output is a 24 kHz mono WAV, watermarked with SynthID.