From the 2 of 5 linked papers with an AI index.
5 papers
WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction
Ke Zhang, Xiaoyang Yu, Haoyu Li +3
WeSep is a modular framework that treats target speaker extraction as a cue‑conditioned learning problem, separating cue modules from the separator backbone to flexibly incorporate…
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Shuai Wang, Zihan Qian, Ke Zhang +9
The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
Jiadong Wang, Ke Zhang, Xinyuan Qian +3
Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Alt…
Interpolating Speaker Identities in Embedding Space for Data Expansion
Tianchi Liu, Ruijie Tao, Qiongqiong Wang +5
The success of deep learning-based speaker verification systems is largely attributed to access to large-scale and diverse speaker identity data. However, collecting data from more…
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
Junjie Li, Ke Zhang, Shuai Wang +3
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world sce…