1 citations · 3 across the 17 of their papers we have counts for
16 papers
BiEAR: A Human Auditory-Inspired Adaptive Binaural Front-end for Multi-Speaker Localisation and Distance Estimation
Hanyu Meng, Eliathamby Ambikairajah, Vidhyasaharan Sethu +2
We present BiEAR, a human auditory-inspired adaptive binaural front-end for multi-speaker localisation and distance estimation. Inspired by medial olivocochlear (MOC) feedback in h…
Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
Jinyang Wu, Zihan Pan, Qiquan Zhang +2
Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for repres…
Analytic Incremental Learning For Sound Source Localization With Imbalance Rectification
Zexia Fan, Yu Chen, Qiquan Zhang +2
Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance…
Fun-Audio-Chat Technical Report
Tongyi Fun Team, Qian Chen, Luyao Cheng +10
Recent advancements in joint speech-text models show great potential for seamless voice interactions. However, existing models face critical challenges: temporal resolution mismatc…
Learning Time-Graph Frequency Representation for Monaural Speech Enhancement
Tingting Wang, Tianrui Wang, Meng Ge +2
The Graph Fourier Transform (GFT) has recently demonstrated promising results in speech enhancement. However, existing GFT-based speech enhancement approaches often employ fixed gr…
Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
Hanyu Meng, Vidhyasaharan Sethu, Eliathamby Ambikairajah +2
In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fix…