MiniMax Speech

Convert text to speech in 40 languages with sound tags.

Playground
APIPricing
~5.97s
Example output
0:00 / 0:00

MiniMax Speech — Text-to-Speech (Text-to-Audio) Model

What is MiniMax Speech?

MiniMax Speech is a text-to-speech API that turns written scripts into natural, expressive spoken audio. It brings the full MiniMax voice family together behind one endpoint: Speech 2.8, Speech 2.6, and Speech 02, each in an HD tier tuned for quality and a Turbo tier tuned for speed. You send text, choose a system voice, and receive ready-to-use audio in formats like mp3, wav, flac, opus, or pcm.

The model is built for vocal authenticity rather than flat, robotic narration. Across its generations, MiniMax Speech has consistently ranked among the top text-to-speech systems on public blind-listening arenas, making it a strong choice when prosody and naturalness matter.

Key Features

  • •Six models, one endpoint: Speech 2.8, 2.6, and 02 in HD and Turbo tiers — switch with a single parameter.
  • •Performed interjections (2.8): inline tags such as (laughs), (sighs), (breath), and (clear-throat) are acted out, not read aloud.
  • •Pause control: insert timed pauses anywhere with the <#x#> marker (seconds).
  • •Emotion control: pick happy, sad, angry, calm, whisper, and more, or let the model auto-select from the text.
  • •40 languages: with language_boost to sharpen recognition of a target language or dialect.
  • •Fine voice shaping: adjust speed (0.5–2), volume (0.1–10), pitch (-12 to +12 semitones), sample rate, and output format.

Best Use Cases

MiniMax Speech fits narration and voiceover, audiobooks, explainer and product videos, e-learning, podcasts, IVR and announcements, and conversational voice agents. HD tiers suit studio-grade narration where clarity and pacing lead; Turbo tiers suit real-time and high-volume synthesis where latency matters. The 2.8 models are the best pick for expressive, conversational scripts thanks to performed sound tags, while 2.6 excels at reading structured text — URLs, phone numbers, dates, and currency — cleanly.

Prompt Tips and Output Quality

Write scripts the way they should be spoken: punctuation drives rhythm. Use <#0.4#>-style markers for deliberate pauses and reserve interjection tags for the 2.8 models. Match language_boost to your script's language for multilingual or code-switched text, and enable text_normalization when a passage is dense with numbers or dates. Keep speed, vol, and pitch near their defaults for the most natural result, then nudge them per voice. For very long scripts, split text into paragraphs for cleaner pacing.