Seed Audio 2.0

Next-gen ByteDance audio model: video dubbing, translation, vocal/SFX/BGM stems, 6-minute output, 30 languages.

Coming Soon
Text To Audio

Seed Audio 2.0 is coming to Segmind

The playground and serverless API for this model aren't live yet. Here's what we know so far — launch details, pricing, and full API docs will appear on this page.

About

Seed Audio 2.0 is the next generation of ByteDance Seed's audio creation model, served to international developers through BytePlus. Where Seed Audio 1.0 turned a text prompt into a complete sound scene of dialogue, music, ambience and effects, 2.0 extends that scene-rendering approach to video: feed it a clip and it generates dubbing and a soundtrack that match what is on screen.

The second generation also changes how you work with the output. Instead of a single mixed track, Seed Audio 2.0 can return independent stems for vocals, sound effects and background music, and it accepts timestamps to place voice-over, effects and music exactly where they belong. Reference capacity doubles and output length triples compared to 1.0, and language coverage grows from a handful of preset-voice languages to 30.

BytePlus has told us to expect Seed Audio 2.0 around 21 September 2026. The provider notes this is not the final version and that feature changes are still possible before release, so the details below may shift.

What to expect

  • Video dubbing: use a video as input to generate dubbing and a soundtrack that match the visuals.
  • Video translation: translate the spoken content into a different language while keeping the original sound effects and background music.
  • Stem separation: independent stem output for vocals, sound effects and background music.
  • Precise control: timestamps to control where human voice-over, sound effects and background music land.
  • Longer output, more references: up to 6 audio references per request, producing up to 6 minutes of audio in a single call.
  • 30 languages: Chinese, English, Japanese, Korean, Mexican Spanish, German, French, Brazilian Portuguese, Thai, Indonesian, Vietnamese, Malay, Arabic (Saudi accent), Castilian Spanish, European Portuguese, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish, Finnish, Danish, Norwegian, Czech, Hungarian, Greek, Romanian and Hindi.
  • Song generation: generate songs with lyrics.

Specifications

ItemSeed Audio 2.0Seed Audio 1.0 (for comparison)
DeveloperByteDance Seed, via BytePlusByteDance Seed, via BytePlus
InputsText, audio references, videoText, audio references, one reference image
Max audio references63
Max output length per request6 minutes2 minutes
OutputMixed track or separate vocal / SFX / BGM stemsSingle mixed track
Languages306 preset-voice languages
Timestamp controlYesNo
Song generation with lyricsYesNo
Expected availabilityAround 21 September 2026Live on Segmind

Availability

BytePlus expects to release Seed Audio 2.0 on or around 21 September 2026. Segmind will add it as soon as the provider opens API access; the playground and API go live on this page, and pricing will be published at that time. Until then, Seed Audio 1.0 is available on Segmind for text-to-audio scene generation with voice cloning.

Sources