Text-to-audio models are increasingly expressive and high-fidelity. You can type in a prompt like "cat meowing" or "glass breaking cinematically" and get a sound effect that often sounds like the request. Re-generate and you get another variation of the same kind of sound. Audio is a very broad category though, and it is unclear what exactly these models capture. Are different models better at different kinds of audio, e.g. natural soundscapes versus music versus sound effects? Do they differ in their range even within the same kind of sound, e.g. a variety of meows versus just a few similar-sounding ones?
As an exploratory first step, we adapted a evaluation method from the procedural content generation (PCG) community, Expressive Range Analysis (ERA), to characterize the distributions of generated audio from various open-weights models across a range of test prompts. The bigger-picture goal is to provide frameworks for systematically analyzing and understanding the capabilities, biases, and boundaries of a given audio model.
Publications:
Funding provided by:
Collaborators (current): Mark Cartwright, Swen Gaudl, Amy K. Hoover, Jonathan Morse