3 papers
cs.SD2025
Recomposer: Event-roll-guided generative audio editing
Daniel P. W. Ellis, Eduardo Fonseca, Ron J. Weiss +7
Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their str…
cs.CL2025
Long-Form Speech Generation with Spoken Language Models
Se Jin Park, Julian Salazar, Aren Jansen +3
We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken l…
cs.LG2024
MELODI: Exploring Memory Compression for Long Contexts
Yinpeng Chen, DeLesley Hutchins, Aren Jansen +3
We present MELODI, a novel memory architecture designed to efficiently process long documents using short context windows. The key principle behind MELODI is to represent short-ter…