Pruna P Video 2

Pruna's next-gen video model: text, image or audio in, 1080p video with native audio — dialogue, music and SFX — out.

Coming Soon
Text To Video

Pruna P Video 2 is coming to Segmind

The playground and serverless API for this model aren't live yet. Here's what we know so far — launch details, pricing, and full API docs will appear on this page.

About

Pruna P Video 2 is the second generation of Pruna AI's video generation model, and it widens the model from a text-to-video generator into a single endpoint that accepts text, an image, or an audio track. The headline change is native audio: P Video 2 generates the soundtrack — dialogue, music and sound effects — together with the picture, instead of leaving you to score the clip afterwards. Pruna also reports stronger lip sync and sharper on-screen text than the first P Video.

The model takes the same inputs as P Video, so prompts and workflows built on the original carry over. Output goes up to 1080p at 24 or 48 fps, with clips of up to 20 seconds. Duration is optional — leave it out and the model picks the clip length from the prompt itself, and when you supply an audio track the video is generated to match the audio's length. A draft mode returns a faster, lower-quality preview for iterating before you commit to a full render.

P Video 2 will be available on Segmind at /v1/p-video-2 with the same single REST call as the rest of the catalog — no GPU infrastructure, no queue management, and billing per second of the video you actually get back.

What to expect

  • Native audio generation — dialogue, music and sound effects produced with the video, with an option to save the video without its audio track.
  • Three ways in — text-to-video, image-to-video (jpg, jpeg, png, webp), and audio-conditioned generation (flac, mp3, wav).
  • Stronger lip sync and sharper rendering of on-screen text compared with P Video.
  • Up to 1080p, at 24 or 48 fps.
  • Clips up to 20 seconds — or omit the duration entirely and let the model choose the length from your prompt.
  • Draft mode for fast, cheaper previews before a full-quality render.
  • Seven aspect ratios — 16:9, 9:16, 4:3, 3:4, 3:2, 2:3 and 1:1 (ignored when you supply an input image).
  • Last-frame conditioning, prompt upsampling, and seeded, reproducible generation.

Specifications

PropertyValue
ProviderPruna AI
ModalitiesText-to-video, image-to-video, audio-conditioned video
Max resolution1080p (720p default)
Frame rate24 or 48 fps
Duration1–20 seconds; optional — the model can choose from the prompt
Aspect ratios16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 1:1
Image input formatsjpg, jpeg, png, webp
Audio input formatsflac, mp3, wav
Draft modeYes — faster, lower-quality preview
Billing basisPer second of returned video, by resolution and draft mode

Availability

Pruna launches P Video 2 on Thursday 10 September 2026. The Segmind endpoint and playground go live on this page shortly after — pricing, parameter reference and example outputs will appear here at launch.

Sources