Showing cs.SDShow all
2 papers · 1 filter
cs.SD2024
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
Junjie Li, Ke Zhang, Shuai Wang +3
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world sce…
cs.SD2024
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
Junjie Li, Ke Zhang, Shuai Wang +3
Deep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms wh…