11 papers
Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books
Karim Benharrak, Oriol Nieto, Bryan Wang +2
Audio dramas weave dialogue, sound effects, and music into immersive stories. Creators often adapt books into audio dramas, but this process remains labor-intensive, requiring them…
AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
William Chen, Prem Seetharaman, Rithesh Kumar +4
Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can h…
Generative Audio Extension and Morphing
Prem Seetharaman, Oriol Nieto, Justin Salamon
In audio-related creative tasks, sound designers often seek to extend and morph different sounds from their libraries. Generative audio models, capable of creating audio using exam…
TAC: Timestamped Audio Captioning
Sonal Kumar, Prem Seetharaman, Ke Chen +8
Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introdu…
Audiocards: Structured Metadata Improves Audio Language Models For Sound Design
Sripathi Sridhar, Prem Seetharaman, Oriol Nieto +2
Sound designers search for sounds in large sound effects libraries using aspects such as sound class or visual context. However, the metadata needed for such search is often missin…
Mix2Morph: Learning Sound Morphing from Noisy Mixes
Annie Chu, Hugo Flores GarcÃa, Oriol Nieto +3
We introduce Mix2Morph, a text-to-audio diffusion model fine-tuned to perform sound morphing without a dedicated dataset of morphs. By finetuning on noisy surrogate mixes at higher…