Sonilo Video to Video — AI Video-to-Music & Sound Effects Scoring
What is Sonilo Video to Video?
Sonilo Video to Video is a video-native AI scoring model that generates an original soundtrack directly from your footage and returns a new MP4 with the audio muxed in. Instead of searching stock libraries or writing genre prompts, it reads the video's pacing, motion, scene changes, and emotional arc, then composes music, places sound effects, or builds a combined music-and-SFX mix that follows the cut. The generated audio matches the exact video length with a natural ending rather than an abrupt loop, so short-form ads, trailers, explainers, product demos, game clips, and social videos feel professionally scored in a single API call.
Key Features
- •Frame-synced, cut-aware scoring generated from the video itself — the prompt is optional.
- •Three modes via
sound_type:music,sfx(frame-accurate Foley and ambience), ormusic_and_sfx. - •
preserve_speechkeeps original dialogue and vocals;duckingautomatically lowers music under speech. - •
sfx_segmentslets you place contiguous, timed sound-effect prompts across the timeline. - •
music_promptandsfx_promptsteer mood, genre, and on-screen action when you want control. - •Returns the original video muxed with a ready-to-publish stereo soundtrack.
Best Use Cases
Use Sonilo Video to Video for short-form content where timing misses are obvious: ad cutdowns, product reveals, drone and travel edits, fitness tutorials, game highlights, and UGC. In testing, music_and_sfx reliably blended cinematic ambient music with matched environmental effects (waves, wind, birds) on a five-second clip, delivering a clean, non-clipping stereo mix. For talking-head clips, vlogs, and interviews, enable preserve_speech so narration stays intact while music sits underneath.
Prompt Tips and Output Quality
Prompt in production language, not just genre terms: describe the emotional arc and the ending, such as build after the intro, leave room for voice-over, or resolve on the final frame. Keep source clips focused and, ideally, video-only. music_and_sfx is the strongest showcase mode because it runs both pipelines. Output audio length tracks the video duration automatically.
FAQs
Does it require a text prompt? No. Prompts are optional; the video drives the score.
Can I keep my narration? Yes — enable preserve_speech (voice only) with ducking.
What input format is best? A clean MP4/MOV, ideally without existing audio.
What does it return? A new MP4 with the generated soundtrack muxed in.
Which mode should I start with? music_and_sfx for a full, mixed result.