4 papers · 1 filter
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
Zhan Jin, Bang Zeng, Peijun Yang +5
Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images,…
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
Bang Zeng, Ming Li
Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spok…
A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound Detection
Yucong Zhang, Juan Liu, Yao Tian +2
In contrast to human speech, machine-generated sounds of the same type often exhibit consistent frequency characteristics and discernible temporal periodicity. However, leveraging…
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
Bang Zeng, Ming Li
Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a refere…