About
Seed Audio 2.0 is the next generation of ByteDance Seed's audio creation model, served to international developers through BytePlus. Where Seed Audio 1.0 turned a text prompt into a complete sound scene of dialogue, music, ambience and effects, 2.0 extends that scene-rendering approach to video: feed it a clip and it generates dubbing and a soundtrack that match what is on screen.
The second generation also changes how you work with the output. Instead of a single mixed track, Seed Audio 2.0 can return independent stems for vocals, sound effects and background music, and it accepts timestamps to place voice-over, effects and music exactly where they belong. Reference capacity doubles and output length triples compared to 1.0, and language coverage grows from a handful of preset-voice languages to 30.
BytePlus has told us to expect Seed Audio 2.0 around 21 September 2026. The provider notes this is not the final version and that feature changes are still possible before release, so the details below may shift.
What to expect
- •Video dubbing: use a video as input to generate dubbing and a soundtrack that match the visuals.
- •Video translation: translate the spoken content into a different language while keeping the original sound effects and background music.
- •Stem separation: independent stem output for vocals, sound effects and background music.
- •Precise control: timestamps to control where human voice-over, sound effects and background music land.
- •Longer output, more references: up to 6 audio references per request, producing up to 6 minutes of audio in a single call.
- •30 languages: Chinese, English, Japanese, Korean, Mexican Spanish, German, French, Brazilian Portuguese, Thai, Indonesian, Vietnamese, Malay, Arabic (Saudi accent), Castilian Spanish, European Portuguese, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish, Finnish, Danish, Norwegian, Czech, Hungarian, Greek, Romanian and Hindi.
- •Song generation: generate songs with lyrics.
Specifications
| Item | Seed Audio 2.0 | Seed Audio 1.0 (for comparison) |
|---|---|---|
| Developer | ByteDance Seed, via BytePlus | ByteDance Seed, via BytePlus |
| Inputs | Text, audio references, video | Text, audio references, one reference image |
| Max audio references | 6 | 3 |
| Max output length per request | 6 minutes | 2 minutes |
| Output | Mixed track or separate vocal / SFX / BGM stems | Single mixed track |
| Languages | 30 | 6 preset-voice languages |
| Timestamp control | Yes | No |
| Song generation with lyrics | Yes | No |
| Expected availability | Around 21 September 2026 | Live on Segmind |
Availability
BytePlus expects to release Seed Audio 2.0 on or around 21 September 2026. Segmind will add it as soon as the provider opens API access; the playground and API go live on this page, and pricing will be published at that time. Until then, Seed Audio 1.0 is available on Segmind for text-to-audio scene generation with voice cloning.
Sources
- •Seed Audio 1.0 product page, ByteDance Seed
- •Introducing the Seed Audio 1.0 Audio Creation Model, ByteDance Seed blog
- •Seed Audio 1.0 API documentation, BytePlus
- •Seed Audio 2.0 capabilities, limits and release window: BytePlus pre-release briefing to Segmind, September 2026