Lyria 3 Pro

Full-length text-to-music songs with vocals and lyrics.

Example output
0:00 / 0:00

Lyria 3 Pro — Text-to-Music Generation Model

What is Lyria 3 Pro?

Lyria 3 Pro is Google DeepMind's premier text-to-music model for generating full-length songs from a prompt. Unlike short-clip models, it composes complete tracks up to roughly three minutes with real structural awareness — distinct intros, verses, choruses, bridges and outros instead of one continuous loop. It reasons through musical structure before generating audio, which keeps energy progression and arrangement coherent from the first note to the last.

Output is high-fidelity 44.1 kHz stereo audio with sung vocals, timed lyrics and full instrumental arrangements. Lyria 3 Pro is also multimodal: alongside your text you can pass up to 10 reference images, and the model composes music inspired by their mood, colour and subject. Every track is embedded with Google's imperceptible SynthID watermark.

Key Features

  • Full songs up to ~3 minutes with prompt-controlled duration
  • Structural control using [Verse], [Chorus], [Bridge], [Intro], [Outro] tags or [0:00 - 0:10] timestamps
  • Custom or automatically generated lyrics, sung with expressive vocals
  • Instrumental-only mode by adding "instrumental only, no vocals"
  • Multilingual lyrics — the language of your prompt sets the vocal language
  • Image-to-music: up to 10 reference images guide mood and style
  • 44.1 kHz high-fidelity stereo output with clean panning and minimal artifacts
  • SynthID watermarking on every generation for provenance

Best Use Cases

Lyria 3 Pro shines wherever polished, production-ready audio matters more than experimental songwriting. In testing, prompts packed with genre, key, BPM, instruments and structure tags returned clean, artifact-free stereo songs with coherent verse and chorus sections. Reach for it to score marketing videos and social content, build game and app soundtracks, produce podcast beds and background music, prototype soundtracks quickly, or turn a mood board of images into a matching track. Because the model is trained on licensed and permissible data and stamps every output with SynthID, it is a strong enterprise-safe choice for commercial work where copyright and provenance are concerns.

Prompt Tips and Output Quality

Be specific: name the genre, instruments, BPM, key and mood in one clear description. Use section tags or timestamps to define the song's arc, and paste your own lyrics when you want exact words — separating lyrics from your musical direction. Add "instrumental only, no vocals" for backing tracks, and state a target length (for example "a 90-second song") to shape duration. Prompt in the language you want the lyrics sung in. In our tests the model faithfully executes a detailed prompt, so prompt quality is the biggest lever on output; vague prompts produce safe, generic results.

FAQs

How long can Lyria 3 Pro tracks be? Up to about three minutes, with duration influenced by your prompt or timestamps.

Does it generate vocals and lyrics? Yes. It sings expressive vocals and can use lyrics you provide or generate them from your prompt, in the prompt's language.

Can I generate instrumental-only music? Yes — add "instrumental only, no vocals" to your prompt for backing tracks and scores.

Can images guide the music? Yes. Provide up to 10 reference images and the model composes music inspired by their mood, colour and subject.

What audio quality does it output? High-fidelity 44.1 kHz stereo, delivered as an audio file with clean stereo separation.

Is the output watermarked? Yes. Every track carries Google's imperceptible SynthID watermark for identifying AI-generated audio.