4 citations · 5 across the 7 of their papers we have counts for
4 papers · 1 filter
USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding
Heng-Jui Chang, Alexander H. Liu, Saurabhchand Bhati +4
Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (S…
ProSE: Diffusion Priors for Speech Enhancement
Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi +4
Speech enhancement (SE) is the foundational task of enhancing the clarity and quality of speech in the presence of non-stationary additive noise. While deterministic deep learning…
Listen2Scene: Interactive material-aware binaural sound propagation for reconstructed 3D scenes
Anton Ratnarajah, Dinesh Manocha
We present an end-to-end binaural audio rendering approach (Listen2Scene) for virtual reality (VR) and augmented reality (AR) applications. We propose a novel neural-network-based…
Improving Reverberant Speech Separation with Multi-stage Training and Curriculum Learning
Rohith Aralikatti, Anton Ratnarajah, Zhenyu Tang +1
We present a novel approach that improves the performance of reverberant speech separation. Our approach is based on an accurate geometric acoustic simulator (GAS) which generates…