3 papers
eess.AS2026
USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding
Heng-Jui Chang, Alexander H. Liu, Saurabhchand Bhati +4
Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (S…
eess.AS2025
ProSE: Diffusion Priors for Speech Enhancement
Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi +4
Speech enhancement (SE) is the foundational task of enhancing the clarity and quality of speech in the presence of non-stationary additive noise. While deterministic deep learning…
cs.SD2024
BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
Anton Ratnarajah, Shi-Xiong Zhang, Dong Yu
We introduce BANC, a neural binaural audio codec designed for efficient speech compression in single and two-speaker scenarios while preserving the spatial location information of…