8 papers
LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization
Zexu Pan, Shengkui Zhao, Yukun Ma +4
Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural au…
E2E-AEC: Implementing an end-to-end neural network learning approach for acoustic echo cancellation
Yiheng Jiang, Biao Tian, Haoxu Wang +4
We propose a novel neural network-based end-to-end acoustic echo cancellation (E2E-AEC) method capable of streaming inference, which operates effectively without reliance on tradit…
FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
Haoxu Wang, Biao Tian, Yiheng Jiang +5
Generative speech enhancement offers a promising alternative to traditional discriminative methods by modeling the distribution of clean speech conditioned on noisy inputs. Post-tr…
Fun-ASR Technical Report
Keyu An, Yanni Chen, Zhigao Chen +35
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep in…
FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
Haoxu Wang, Yiheng Jiang, Gang Qiao +2
Speech separation always faces the challenge of handling prolonged time sequences. Past methods try to reduce sequence lengths and use the Transformer to capture global information…
Exploring Efficient Directional and Distance Cues for Regional Speech Separation
Yiheng Jiang, Haoxu Wang, Yafeng Chen +2
In this paper, we introduce a neural network-based method for regional speech separation using a microphone array. This approach leverages novel spatial cues to extract the sound s…