2 papers
cs.SD2025
Expressive Range Characterization of Open Text-to-Audio Models
Jonathan Morse, Azadeh Naderi, Swen Gaudl +3
Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional…
cs.SD2024
EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation
Mithun Manivannan, Vignesh Nethrapalli, Mark Cartwright
Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. Howev…