activity
20242026
most citedFun-Audio-Chat Technical Report

1 citations · 3 across the 17 of their papers we have counts for

collaborators

16 papers

eess.AS2026

BiEAR: A Human Auditory-Inspired Adaptive Binaural Front-end for Multi-Speaker Localisation and Distance Estimation

Hanyu Meng, Eliathamby Ambikairajah, Vidhyasaharan Sethu +2

We present BiEAR, a human auditory-inspired adaptive binaural front-end for multi-speaker localisation and distance estimation. Inspired by medial olivocochlear (MOC) feedback in h…

cs.SD2026

Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection

Jinyang Wu, Zihan Pan, Qiquan Zhang +2

Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for repres…

cs.SD2026

Analytic Incremental Learning For Sound Source Localization With Imbalance Rectification

Zexia Fan, Yu Chen, Qiquan Zhang +2

Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance…

cs.CL20261 cited

Fun-Audio-Chat Technical Report

Tongyi Fun Team, Qian Chen, Luyao Cheng +10

Recent advancements in joint speech-text models show great potential for seamless voice interactions. However, existing models face critical challenges: temporal resolution mismatc…

eess.AS2025

Learning Time-Graph Frequency Representation for Monaural Speech Enhancement

Tingting Wang, Tianrui Wang, Meng Ge +2

The Graph Fourier Transform (GFT) has recently demonstrated promising results in speech enhancement. However, existing GFT-based speech enhancement approaches often employ fixed gr…

eess.AS2025

Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing

Hanyu Meng, Vidhyasaharan Sethu, Eliathamby Ambikairajah +2

In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fix…