Lyria 3

Generate 30-second songs with vocals from text or images.

Example output
0:00 / 0:00

Lyria 3 — Text-to-Music AI Model

What is Lyria 3?

Lyria 3 is Google DeepMind's music generation model, served on Segmind as a text-to-audio API. From a single text prompt, it composes a 30-second, 44.1 kHz stereo clip complete with vocals, timed lyrics, and full instrumental arrangements — or an instrumental-only backing track when you ask for one. Lyria 3 also accepts up to 10 reference images, so you can turn a photo's mood, color, and subject into a matching soundtrack. Before it renders audio, the model reasons through musical structure (intro, verse, chorus, bridge) to keep the composition coherent from the first note to the last.

Key Features

  • Text-to-music and image-to-music generation from natural language prompts
  • 30-second, 44.1 kHz high-fidelity stereo MP3 output
  • Vocals with time-aligned lyrics, or instrumental-only tracks on request
  • Multilingual lyrics — write your prompt in the language you want to hear
  • Structure control with [Verse], [Chorus], [Bridge] tags and [0:00 - 0:10] timestamps
  • Every track carries an imperceptible SynthID watermark for AI transparency

Best Use Cases

Lyria 3 is built for creators who need custom, royalty-aware audio fast: social and short-form video soundtracks, background music for games and apps, marketing jingles, podcast intros, lo-fi study loops, and demo songs with sung hooks. The image-to-music workflow is ideal for auto-scoring campaign assets or matching a track to a brand photo. In testing, a single prompt reliably produced a full-band, radio-ready mix with clean lead vocals in about 30 seconds.

Prompt Tips and Output Quality

Be specific: name the genre, instruments, BPM, key, and mood. Use section tags or timestamps to shape progression, and paste your own lyrics for sung vocals. Add "instrumental only, no vocals" for a clean backing track. Vague prompts yield generic results, so layer detail. Results vary between calls since generation is non-deterministic.

FAQs

Does Lyria 3 generate vocals and lyrics? Yes — it sings time-aligned lyrics and can also produce instrumental-only tracks.

How long are the clips? Each generation is a fixed 30-second, 44.1 kHz stereo MP3.

Can I generate music from an image? Yes — supply up to 10 reference images to guide mood and style.

Does it support other languages? Yes — lyrics are generated in the language of your prompt.

Are outputs watermarked? Yes — every track includes an imperceptible SynthID watermark.

Can I request longer, full-length songs? For multi-minute tracks with detailed structure, use Lyria 3 Pro.