5 papers
SUNAC: Source-aware Unified Neural Audio Codec
Ryo Aihara, Yoshiki Masuyama, Francesco Paissan +3
Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures…
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
Ryo Aihara, Yoshiki Masuyama, Gordon Wichern +2
Neural audio codecs (NACs), which use neural networks to generate compact audio representations, have garnered interest for their applicability to many downstream tasks -- especial…
Direction-Aware Neural Acoustic Fields for Few-Shot Interpolation of Ambisonic Impulse Responses
Christopher Ick, Gordon Wichern, Yoshiki Masuyama +2
The characteristics of a sound field are intrinsically linked to the geometric and spatial properties of the environment surrounding a sound source and a listener. The physics of s…
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang +3
Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a…
Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
Christopher Ick, Gordon Wichern, Yoshiki Masuyama +2
This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1)…