collaborators

11 papers

eess.AS2026

LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization

Zexu Pan, Shengkui Zhao, Yukun Ma +4

Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural au…

eess.AS2026

Beyond Lips: Integrating Gesture and Lip Cues for Robust Audio-visual Speaker Extraction

Zexu Pan, Xinyuan Qian, Shengkui Zhao +2

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human com…

cs.SD2026

E2E-AEC: Implementing an end-to-end neural network learning approach for acoustic echo cancellation

Yiheng Jiang, Biao Tian, Haoxu Wang +4

We propose a novel neural network-based end-to-end acoustic echo cancellation (E2E-AEC) method capable of streaming inference, which operates effectively without reliance on tradit…

eess.AS2026

FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning

Haoxu Wang, Biao Tian, Yiheng Jiang +5

Generative speech enhancement offers a promising alternative to traditional discriminative methods by modeling the distribution of clean speech conditioned on noisy inputs. Post-tr…

cs.CL2025

Fun-ASR Technical Report

Keyu An, Yanni Chen, Zhigao Chen +35

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep in…

cs.SD2025

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Shengkui Zhao, Zexu Pan, Bin Ma

This paper introduces ClearerVoice-Studio, an open-source, AI-powered speech processing toolkit designed to bridge cutting-edge research and practical application. Unlike broad pla…