2 papers
eess.AS2026
Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
Fuyuan Feng, Wenbin Zhang, Yu Gao +3
Personal Voice Activity Detection (PVAD) is crucial for identifying target speaker segments in the mixture, yet its performance heavily depends on the quality of speaker embeddings…
eess.AS2025
End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
Kangqi Jing, Wenbin Zhang, Yu Gao
Target Speaker Extraction (TSE) plays a critical role in enhancing speech signals in noisy and multi-speaker environments. This paper presents an end-to-end TSE model that incorpor…