target speaker extraction 2audio separation 1cue composability 1evaluation metrics 1modular framework 1multimodal cues 1online streaming 1real-world recordings 1speech separation 1
From the 2 of 9 linked papers with an AI index.
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction
Ke Zhang, Xiaoyang Yu, Haoyu Li +3
WeSep is a modular framework that treats target speaker extraction as a cue‑conditioned learning problem, separating cue modules from the separator backbone to flexibly incorporate…
eess.AS2026
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Shuai Wang, Zihan Qian, Ke Zhang +9
The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…
eess.AS2024
Multi-Level Speaker Representation for Target Speaker Extraction
Ke Zhang, Junjie Li, Shuai Wang +4
Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the refere…