MiniMax Speech Serverless API

Convert text to speech in 40 languages with sound tags.

PlaygroundAPI
Pricing
~5.97s
POST /v2/minimax-speech · submit + poll
 1# pip install "segmind>=1.1.0"
 2# export SEGMIND_API_KEY="YOUR_API_KEY"
 3import segmind
 4
 5# Async (v2): submit to the queue and block until COMPLETED.
 6# run() returns the final result dict (600s deadline, 1.0s poll by default).
 7result = segmind.run(
 8    "minimax-speech",
 9    text="The saffron goes in now — (sniffs) oh, that aroma. <#0.4#> We rest the paella for exactly twelve minutes, no peeking. (laughs) I know, the hardest part. (breath) So... worth the wait? One bite of Valencia, and absolutely — yes.",
10    model="speech-2.8-hd",
11    voice_id="English_expressive_narrator",
12    language_boost="auto",
13    speed=1,
14    vol=1,
15    pitch=0,
16    format="mp3",
17    sample_rate=44100,
18    text_normalization=False,
19)
20print(result["status"])                      # COMPLETED
21print(result.get("output"))                  # model output (e.g. media URL)
22print(result["metrics"]["inference_time"])   # server compute seconds
23
24# --- Or submit + poll manually (track request_id, control the cadence) ---
25from segmind import SegmindClient, InferenceFailed, InferenceTimeout
26
27client = SegmindClient()                      # reads SEGMIND_API_KEY
28payload = {
29    "text": "The saffron goes in now — (sniffs) oh, that aroma. <#0.4#> We rest the paella for exactly twelve minutes, no peeking. (laughs) I know, the hardest part. (breath) So... worth the wait? One bite of Valencia, and absolutely — yes.",
30    "model": "speech-2.8-hd",
31    "voice_id": "English_expressive_narrator",
32    "language_boost": "auto",
33    "speed": 1,
34    "vol": 1,
35    "pitch": 0,
36    "format": "mp3",
37    "sample_rate": 44100,
38    "text_normalization": False,
39}
40job = client.submit_async("minimax-speech", **payload)
41print(job.request_id)                         # available immediately
42try:
43    result = job.wait(timeout=600, interval=1.0)
44except InferenceTimeout as e:
45    print("still running:", e.request_id)
46except InferenceFailed as e:
47    print("failed:", e.detail)

API Endpoint

POSThttps://api.segmind.com/v1/minimax-speech

Parameters

textrequired
string

Text to synthesize, up to 10,000 characters. Insert pauses with <#x#> (seconds, e.g. <#0.5#>); on the 2.8 models, interjection tags such as (laughs), (sighs) or (breath) are performed rather than read.

emotionoptional
string

Leave unset to let the model pick the most natural emotion from the text.

Allowed values :
Happy→"happy"
Sad→"sad"
Angry→"angry"
Fearful→"fearful"
Disgusted→"disgusted"
Surprised→"surprised"
Calm→"calm"
Fluent→"fluent"
Whisper→"whisper"
formatoptional
string

Audio output format: mp3, wav, flac, opus or pcm. Use mp3 for web, wav or flac for editing.

Default: "mp3"
Allowed values :
"mp3""wav""flac""opus""pcm"
language_boostoptional
string

Improves recognition of the given language or dialect; auto detects it.

Allowed values (41 total):
"auto""Chinese""Chinese,Yue""English""Arabic""Russian""Spanish""French""Portuguese""German"+31 more
modeloptional
string

MiniMax speech model. HD favours quality, Turbo favours speed and costs less. whisper emotion is not available on 2.8; Persian, Filipino and Tamil need 2.6 or 2.8.

Default: "speech-2.8-hd"
Allowed values :
Speech 2.8 HD→"speech-2.8-hd"
Speech 2.8 Turbo→"speech-2.8-turbo"
Speech 2.6 HD→"speech-2.6-hd"
Speech 2.6 Turbo→"speech-2.6-turbo"
Speech 02 HD→"speech-02-hd"
Speech 02 Turbo→"speech-02-turbo"
pitchoptional
integer

Semitone shift.

Default: 0Range: -12 - 12
sample_rateoptional
integer

Sample rate in Hz, 8000 to 44100. Use 44100 for studio-grade audio, 16000 for telephony.

Default: 32000
Allowed values :
8000 Hz→8000
16000 Hz→16000
22050 Hz→22050
24000 Hz→24000
32000 Hz→32000
44100 Hz→44100
speedoptional
number

Speech rate, 0.5 to 2. Lower for audiobooks, raise for brisk announcements; keep near 1 for natural delivery.

Default: 1Range: 0.5 - 2
text_normalizationoptional
boolean

Better reading of digits and dates in Chinese and English, slightly slower.

Default: false
voice_idoptional
string

A MiniMax system voice id, e.g. English_expressive_narrator, English_Graceful_Lady, English_Persuasive_Man, Japanese_Whisper_Belle, Cantonese_GentleLady (set language_boost to Chinese,Yue).

Default: "English_expressive_narrator"
voloptional
number

Output loudness, 0.1 to 10. Keep at 1 for most cases; raise slightly to lift a quiet voice.

Default: 1Range: 0.1 - 10

Response Type

Returns: Audio

Asynchronous requests (v2)

Use Async for video, long-running (>~60s), or high-concurrency workloads; Sync is simplest for fast image & LLM calls. Async submits a request and you poll it to completion.

  1. 1
    POST /v2/minimax-speech

    Submit — returns request_id, status_url, response_url

  2. 2
    GET /v2/requests/{id}/status

    Poll — until COMPLETED or FAILED

  3. 3
    GET /v2/requests/{id}

    Result — final response body

Status states

QUEUED— Accepted, waiting for a worker
PROCESSING— Running on a worker
COMPLETED— Done — result body is ready
FAILED— Errored (incl. content/RAI blocks)
  • A FAILED request is served as HTTP 422 — the body still carries the error detail.
  • An unknown or expired request_id returns HTTP 404.
  • Results are retained for 1 hour, then expire.
  • Content / RAI blocks surface as FAILED, not a separate state.
  • Track completion by polling the status endpoint.

Common Error Codes

The API returns standard HTTP status codes. Detailed error messages are provided in the response body.

400

Bad Request

Invalid parameters or request format

401

Unauthorized

Missing or invalid API key

403

Forbidden

Insufficient permissions

404

Not Found

Model or endpoint not found

406

Insufficient Credits

Not enough credits to process request

429

Rate Limited

Too many requests

500

Server Error

Internal server error

502

Bad Gateway

Service temporarily unavailable

504

Timeout

Request timed out