13 papers
WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction
Ke Zhang, Xiaoyang Yu, Haoyu Li +3
WeSep is a modular framework that treats target speaker extraction as a cue‑conditioned learning problem, separating cue modules from the separator backbone to flexibly incorporate…
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Shuai Wang, Zihan Qian, Ke Zhang +9
The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…
Detect, Attend and Extract: Keyword Guided Target Speaker Extraction
Haoyu Li, Yu Xi, Yidi Jiang +5
Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on…
A Unified and Reproducible Experimentation Framework for Speech Understanding
Jing Peng, Junhao Du, Chenghao Wang +21
Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable evaluations caused by mismatched…
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
Jing Peng, Ziyi Chen, Haoyu Li +7
We study timestamped speaker-attributed automatic speech recognition (SA-ASR) for long-form, multi-party speech with overlap. In this setting, chunk-wise inference must preserve me…
Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
Haoyu Li, Mingyang Han, Yu Xi +8
Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ab…