6 papers · 1 filter
LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization
Zexu Pan, Shengkui Zhao, Yukun Ma +4
Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural au…
Beyond Lips: Integrating Gesture and Lip Cues for Robust Audio-visual Speaker Extraction
Zexu Pan, Xinyuan Qian, Shengkui Zhao +2
Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human com…
FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
Haoxu Wang, Biao Tian, Yiheng Jiang +5
Generative speech enhancement offers a promising alternative to traditional discriminative methods by modeling the distribution of clean speech conditioned on noisy inputs. Post-tr…
Online Audio-Visual Autoregressive Speaker Extraction
Zexu Pan, Wupeng Wang, Shengkui Zhao +4
This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less ex…
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
Zexu Pan, Shengkui Zhao, Tingting Wang +4
Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. Ho…
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
Kun Zhou, You Zhang, Dianwen Ng +3
Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of e…