11 papers
Diagnosing Neural Convergence with Topological Alignment Spectra
Tiago F. Tavares, Fabio Ayres, Paris Smaragdis
Representational similarity in neural networks is inherently scale-dependent, yet widely used metrics such as Centered Kernel Alignment (CKA) and Procrustes analysis provide only g…
Gencho: Room Impulse Response Generation from Reverberant Speech and Text via Diffusion Transformers
Jackie Lin, Jiaqi Su, Nishit Anand +3
Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability a…
PromptSep: Generative Audio Separation via Multimodal Prompting
Yutong Wen, Ke Chen, Prem Seetharaman +7
Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based…
Re-Bottleneck: Latent Re-Structuring for Neural Audio Autoencoders
Dimitrios Bralios, Jonah Casebeer, Paris Smaragdis
Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitatio…
Learning to Upsample and Upmix Audio in the Latent Domain
Dimitrios Bralios, Paris Smaragdis, Jonah Casebeer
Neural audio autoencoders create compact latent representations that preserve perceptually important information, serving as the foundation for both modern audio compression system…
Combolutional Neural Networks
Cameron Churchwell, Minje Kim, Paris Smaragdis
Selecting appropriate inductive biases is an essential step in the design of machine learning models, especially when working with audio, where even short clips may contain million…