works on

From the 2 of 9 linked papers with an AI index.

activity
20242026
collaborators

9 papers

eess.AS2026

WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction

Ke Zhang, Xiaoyang Yu, Haoyu Li +3

WeSep is a modular framework that treats target speaker extraction as a cue‑conditioned learning problem, separating cue modules from the separator backbone to flexibly incorporate…

eess.AS2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Shuai Wang, Zihan Qian, Ke Zhang +9

The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…

cs.SD2026

MeMo: Attentional Momentum for Real-Time Audio-Visual Target Speaker Extraction Under Impaired Visual Conditions

Junjie Li, Wenxuan Wu, Shuai Wang +4

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the perform…

cs.SD2025

ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction

Wenxuan Wu, Shuai Wang, Xixin Wu +2

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic…

cs.SD2025

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

Wenxuan Wu, Shuai Wang, Xixin Wu +2

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguist…

cs.SD2025

AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction

Wenxuan Wu, Xueyuan Chen, Shuai Wang +5

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recent…