Lyria 3

Generate 30-second songs with vocals from text or images.

Playground
APIPricing
~18.06s
Example output
0:00 / 0:00

Lyria 3 — Text-to-Music AI Model

What is Lyria 3?

Lyria 3 is Google DeepMind's music generation model, served on Segmind as a text-to-audio API. From a single text prompt, it composes a 30-second, 44.1 kHz stereo clip complete with vocals, timed lyrics, and full instrumental arrangements — or an instrumental-only backing track when you ask for one. Lyria 3 also accepts up to 10 reference images, so you can turn a photo's mood, color, and subject into a matching soundtrack. Before it renders audio, the model reasons through musical structure (intro, verse, chorus, bridge) to keep the composition coherent from the first note to the last.

Key Features

  • •Text-to-music and image-to-music generation from natural language prompts
  • •30-second, 44.1 kHz high-fidelity stereo MP3 output
  • •Vocals with time-aligned lyrics, or instrumental-only tracks on request
  • •Multilingual lyrics — write your prompt in the language you want to hear
  • •Structure control with [Verse], [Chorus], [Bridge] tags and [0:00 - 0:10] timestamps
  • •Every track carries an imperceptible SynthID watermark for AI transparency

Best Use Cases

Lyria 3 is built for creators who need custom, royalty-aware audio fast: social and short-form video soundtracks, background music for games and apps, marketing jingles, podcast intros, lo-fi study loops, and demo songs with sung hooks. The image-to-music workflow is ideal for auto-scoring campaign assets or matching a track to a brand photo. In testing, a single prompt reliably produced a full-band, radio-ready mix with clean lead vocals in about 30 seconds.

Prompt Tips and Output Quality

Be specific: name the genre, instruments, BPM, key, and mood. Use section tags or timestamps to shape progression, and paste your own lyrics for sung vocals. Add "instrumental only, no vocals" for a clean backing track. Vague prompts yield generic results, so layer detail. Results vary between calls since generation is non-deterministic.