🎭 ScenA — Reference-Driven Multi-Speaker Audio Scenes

Generate a whole audio scene — dialogue, paralinguistics, sound effects and ambience — from a text prompt, with the speakers' voices set by one to three reference clips.

model · code · project page

Prompting — refer to speakers as "the speaker from reference 1", "reference 2", … matching the order of the clips below. Put spoken words in quotes; describe sound effects and ambience in plain prose. Reference clips work best as clean single-speaker speech up to ~20 s.

2 20
Examples — prompts and reference voices from the ScenA project page