Seed Audio 1.5

Next ByteDance audio model: video dubbing, translation, vocal/SFX/BGM layers, 6-minute output, 30 languages.

Coming Soon
Text To Audio

Seed Audio 1.5 is coming to Segmind

The playground and serverless API for this model aren't live yet. Here's what we know so far — launch details, pricing, and full API docs will appear on this page.

About

Seed Audio 1.5 is the next release of ByteDance Seed's audio creation model, served to international developers through BytePlus. Where Seed Audio 1.0 turned a text prompt into a complete sound scene of dialogue, music, ambience and effects, 1.5 extends that scene-rendering approach to video: feed it a clip and it generates dubbing and a soundtrack that match what is on screen, with full context carried across the whole video rather than scene by scene.

The release also changes how you work with the output. Instead of a single mixed track, Seed Audio 1.5 can return independent layers for vocals, sound effects and background music, so each voice and the score can be placed on their own timeline tracks. Timestamps let you control exactly where voice-over, effects and music land. Reference capacity doubles and output length triples compared to 1.0, and consistency holds across the full run, with no jumps at 30-second marks or sudden music shifts. Language coverage grows from a handful of preset-voice languages to 30, including Hindi.

BytePlus's pre-release brief pointed to a 21 September 2026 release. The provider has noted this is not the final version and that feature changes are still possible, so the details below may shift.

What to expect

  • Video dubbing: use a video as input to generate dubbing and a soundtrack that match the visuals, with context maintained across the entire video.
  • Video translation: translate the spoken content into a different language while keeping the original sound effects and background music.
  • Layer separation: independent stem output for vocals, sound effects and background music, with separate tracks per voice.
  • Precise control: timestamps to control where human voice-over, sound effects and background music land.
  • Longer output, more references: up to 6 audio references per request, producing up to 6 minutes of consistent audio in a single call.
  • 30 languages: Chinese, English, Japanese, Korean, Mexican Spanish, German, French, Brazilian Portuguese, Thai, Indonesian, Vietnamese, Malay, Arabic (Saudi accent), Castilian Spanish, European Portuguese, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish, Finnish, Danish, Norwegian, Czech, Hungarian, Greek, Romanian and Hindi.
  • Song generation: generate songs with lyrics.

Specifications

ItemSeed Audio 1.5Seed Audio 1.0 (for comparison)
DeveloperByteDance Seed, via BytePlusByteDance Seed, via BytePlus
InputsText, audio references, videoText, audio references, one reference image
Max audio references63
Max output length per request6 minutes2 minutes
OutputMixed track or separate vocal / SFX / BGM layersSingle mixed track
Languages306 preset-voice languages
Timestamp controlYesNo
Song generation with lyricsYesNo
Expected availabilityPre-release; brief pointed to 21 September 2026Live on Segmind

Availability

BytePlus's pre-release brief pointed to 21 September 2026. Segmind will add Seed Audio 1.5 as soon as the provider opens API access; the playground and API go live on this page, and pricing will be published at that time. Until then, Seed Audio 1.0 is available on Segmind for text-to-audio scene generation with voice cloning.

Sources