2 papers
cs.SD2026
Training-Free Multi-Step Inference for Target Speaker Extraction
Zhenghai You, Ying Shi, Lantian Li +1
Target speaker extraction (TSE) aims to recover a target speaker's speech from a mixture using a reference utterance as a cue. Most TSE systems adopt conditional auto-encoder archi…
cs.SD2025
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
Zhenghai You, Zhenyu Zhou, Lantian Li +1
Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is la…