🎠ScenA — Reference-Driven Multi-Speaker Audio Scenes
Generate a whole audio scene — dialogue, paralinguistics, sound effects and ambience — from a text prompt, with the speakers' voices set by one to three reference clips.
model · code · project page
Prompting — refer to speakers as "the speaker from reference 1", "reference 2", … matching the order of the clips below. Put spoken words in quotes; describe sound effects and ambience in plain prose. Reference clips work best as clean single-speaker speech up to ~20 s.
2 20
Examples — prompts and reference voices from the ScenA project page