2 papers
eess.AS2025
Cross-attention Inspired Selective State Space Models for Target Sound Extraction
Donghang Wu, Yiwen Wang, Xihong Wu +1
The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given cl…
cs.SD2024
DENSE: Dynamic Embedding Causal Target Speech Extraction
Yiwen Wang, Zeyu Yuan, Xihong Wu
Target speech extraction (TSE) focuses on extracting the speech of a specific target speaker from a mixture of signals. Existing TSE models typically utilize static embeddings as c…