3 papers
cs.SD2026
A Generative-First Neural Audio Autoencoder
Jonah Casebeer, Ge Zhu, Zhepei Wang +1
Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single…
cs.SD2026
TAC: Timestamped Audio Captioning
Sonal Kumar, Prem Seetharaman, Ke Chen +8
Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introdu…
cs.SD2026
Rethinking Music Captioning with Music Metadata LLMs
Irmak Bukey, Zhepei Wang, Chris Donahue +1
Music captioning, or the task of generating a natural language description of music, is useful for both music understanding and controllable music generation. Training captioning m…