activity
20242026
collaborators

6 papers

cs.SD2026

SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios

Ziyang Jiang, Yu Chen, Zexu Pan +5

Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-b…

cs.SD2026

MeMo: Attentional Momentum for Real-Time Audio-Visual Target Speaker Extraction Under Impaired Visual Conditions

Junjie Li, Wenxuan Wu, Shuai Wang +4

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the perform…

cs.SD2026

pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues

Ziyang Jiang, Jiahe Lei, Xueyan Chen +4

Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has…

cs.SD2025

Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation

Wupeng Wang, Zexu Pan, Xinke Li +2

Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a pro…

cs.SD2025

Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation

Wupeng Wang, Zexu Pan, Jingru Lin +2

Speech separation seeks to isolate individual speech signals from a multi-talk speech mixture. Despite much progress, a system well-trained on synthetic data often experiences perf…

cs.SD2024

Speech Separation with Pretrained Frontend to Minimize Domain Mismatch

Wupeng Wang, Zexu Pan, Xinke Li +2

Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of ta…