MiniMax Speech Serverless API
Convert text to speech in 40 languages with sound tags.
POST /v2/minimax-speech · submit + poll 1# pip install "segmind>=1.1.0"
2# export SEGMIND_API_KEY="YOUR_API_KEY"
3import segmind
4
5# Async (v2): submit to the queue and block until COMPLETED.
6# run() returns the final result dict (600s deadline, 1.0s poll by default).
7result = segmind.run(
8 "minimax-speech",
9 text="The saffron goes in now — (sniffs) oh, that aroma. <#0.4#> We rest the paella for exactly twelve minutes, no peeking. (laughs) I know, the hardest part. (breath) So... worth the wait? One bite of Valencia, and absolutely — yes.",
10 model="speech-2.8-hd",
11 voice_id="English_expressive_narrator",
12 language_boost="auto",
13 speed=1,
14 vol=1,
15 pitch=0,
16 format="mp3",
17 sample_rate=44100,
18 text_normalization=False,
19)
20print(result["status"]) # COMPLETED
21print(result.get("output")) # model output (e.g. media URL)
22print(result["metrics"]["inference_time"]) # server compute seconds
23
24# --- Or submit + poll manually (track request_id, control the cadence) ---
25from segmind import SegmindClient, InferenceFailed, InferenceTimeout
26
27client = SegmindClient() # reads SEGMIND_API_KEY
28payload = {
29 "text": "The saffron goes in now — (sniffs) oh, that aroma. <#0.4#> We rest the paella for exactly twelve minutes, no peeking. (laughs) I know, the hardest part. (breath) So... worth the wait? One bite of Valencia, and absolutely — yes.",
30 "model": "speech-2.8-hd",
31 "voice_id": "English_expressive_narrator",
32 "language_boost": "auto",
33 "speed": 1,
34 "vol": 1,
35 "pitch": 0,
36 "format": "mp3",
37 "sample_rate": 44100,
38 "text_normalization": False,
39}
40job = client.submit_async("minimax-speech", **payload)
41print(job.request_id) # available immediately
42try:
43 result = job.wait(timeout=600, interval=1.0)
44except InferenceTimeout as e:
45 print("still running:", e.request_id)
46except InferenceFailed as e:
47 print("failed:", e.detail) 1# pip install "segmind>=1.1.0"
2# export SEGMIND_API_KEY="YOUR_API_KEY"
3import segmind
4
5# Async (v2): submit to the queue and block until COMPLETED.
6# run() returns the final result dict (600s deadline, 1.0s poll by default).
7result = segmind.run(
8 "minimax-speech",
9 text="The saffron goes in now — (sniffs) oh, that aroma. <#0.4#> We rest the paella for exactly twelve minutes, no peeking. (laughs) I know, the hardest part. (breath) So... worth the wait? One bite of Valencia, and absolutely — yes.",
10 model="speech-2.8-hd",
11 voice_id="English_expressive_narrator",
12 language_boost="auto",
13 speed=1,
14 vol=1,
15 pitch=0,
16 format="mp3",
17 sample_rate=44100,
18 text_normalization=False,
19)
20print(result["status"]) # COMPLETED
21print(result.get("output")) # model output (e.g. media URL)
22print(result["metrics"]["inference_time"]) # server compute seconds
23
24# --- Or submit + poll manually (track request_id, control the cadence) ---
25from segmind import SegmindClient, InferenceFailed, InferenceTimeout
26
27client = SegmindClient() # reads SEGMIND_API_KEY
28payload = {
29 "text": "The saffron goes in now — (sniffs) oh, that aroma. <#0.4#> We rest the paella for exactly twelve minutes, no peeking. (laughs) I know, the hardest part. (breath) So... worth the wait? One bite of Valencia, and absolutely — yes.",
30 "model": "speech-2.8-hd",
31 "voice_id": "English_expressive_narrator",
32 "language_boost": "auto",
33 "speed": 1,
34 "vol": 1,
35 "pitch": 0,
36 "format": "mp3",
37 "sample_rate": 44100,
38 "text_normalization": False,
39}
40job = client.submit_async("minimax-speech", **payload)
41print(job.request_id) # available immediately
42try:
43 result = job.wait(timeout=600, interval=1.0)
44except InferenceTimeout as e:
45 print("still running:", e.request_id)
46except InferenceFailed as e:
47 print("failed:", e.detail)API Endpoint
https://api.segmind.com/v1/minimax-speechParameters
textrequiredstringText to synthesize, up to 10,000 characters. Insert pauses with <#x#> (seconds, e.g. <#0.5#>); on the 2.8 models, interjection tags such as (laughs), (sighs) or (breath) are performed rather than read.
emotionoptionalstringLeave unset to let the model pick the most natural emotion from the text.
"happy""sad""angry""fearful""disgusted""surprised""calm""fluent""whisper"formatoptionalstringAudio output format: mp3, wav, flac, opus or pcm. Use mp3 for web, wav or flac for editing.
"mp3""mp3""wav""flac""opus""pcm"language_boostoptionalstringImproves recognition of the given language or dialect; auto detects it.
"auto""Chinese""Chinese,Yue""English""Arabic""Russian""Spanish""French""Portuguese""German"+31 moremodeloptionalstringMiniMax speech model. HD favours quality, Turbo favours speed and costs less. whisper emotion is not available on 2.8; Persian, Filipino and Tamil need 2.6 or 2.8.
"speech-2.8-hd""speech-2.8-hd""speech-2.8-turbo""speech-2.6-hd""speech-2.6-turbo""speech-02-hd""speech-02-turbo"pitchoptionalintegerSemitone shift.
0Range: -12 - 12sample_rateoptionalintegerSample rate in Hz, 8000 to 44100. Use 44100 for studio-grade audio, 16000 for telephony.
3200080001600022050240003200044100speedoptionalnumberSpeech rate, 0.5 to 2. Lower for audiobooks, raise for brisk announcements; keep near 1 for natural delivery.
1Range: 0.5 - 2text_normalizationoptionalbooleanBetter reading of digits and dates in Chinese and English, slightly slower.
falsevoice_idoptionalstringA MiniMax system voice id, e.g. English_expressive_narrator, English_Graceful_Lady, English_Persuasive_Man, Japanese_Whisper_Belle, Cantonese_GentleLady (set language_boost to Chinese,Yue).
"English_expressive_narrator"voloptionalnumberOutput loudness, 0.1 to 10. Keep at 1 for most cases; raise slightly to lift a quiet voice.
1Range: 0.1 - 10Response Type
Returns: Audio
Asynchronous requests (v2)
Use Async for video, long-running (>~60s), or high-concurrency workloads; Sync is simplest for fast image & LLM calls. Async submits a request and you poll it to completion.
- 1
POST /v2/minimax-speechSubmit — returns request_id, status_url, response_url
- 2
GET /v2/requests/{id}/statusPoll — until COMPLETED or FAILED
- 3
GET /v2/requests/{id}Result — final response body
Status states
- A FAILED request is served as HTTP 422 — the body still carries the error detail.
- An unknown or expired request_id returns HTTP 404.
- Results are retained for 1 hour, then expire.
- Content / RAI blocks surface as FAILED, not a separate state.
- Track completion by polling the status endpoint.
Common Error Codes
The API returns standard HTTP status codes. Detailed error messages are provided in the response body.
Bad Request
Invalid parameters or request format
Unauthorized
Missing or invalid API key
Forbidden
Insufficient permissions
Not Found
Model or endpoint not found
Insufficient Credits
Not enough credits to process request
Rate Limited
Too many requests
Server Error
Internal server error
Bad Gateway
Service temporarily unavailable
Timeout
Request timed out