collaborators

5 papers

cs.SD2026

SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios

Ziyang Jiang, Yu Chen, Zexu Pan +5

Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-b…

cs.SD2026

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

Junjie Li, Wenxuan Wu, Shuai Wang +4

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the perform…

cs.SD2026

pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues

Ziyang Jiang, Jiahe Lei, Xueyan Chen +4

Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has…

cs.SD2025

Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation

Wupeng Wang, Zexu Pan, Xinke Li +2

Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a pro…

cs.SD2025

Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation

Wupeng Wang, Zexu Pan, Jingru Lin +2

Speech separation seeks to isolate individual speech signals from a multi-talk speech mixture. Despite much progress, a system well-trained on synthetic data often experiences perf…