collaborators

5 papers

eess.AS2026

Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models

Bang Zeng, Beilong Tang, Wang Xiang +1

Target speaker extraction (TSE) aims to recover the speech of a desired speaker from a mixture given a short enrollment utterance, while speech enhancement (SE) focuses on improvin…

eess.AS2026

Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion

Zhan Jin, Bang Zeng, Peijun Yang +5

Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images,…

cs.LG2025

LauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language Models

Beilong Tang, Bang Zeng, Ming Li

We propose LauraTSE, an Auto-Regressive Decoder-Only Language Model for Target Speaker Extraction built upon the LauraGPT backbone. LauraTSE employs a small-scale auto-regressive d…

eess.AS2025

USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

Bang Zeng, Ming Li

Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a refere…

eess.AS2025

Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection

Bang Zeng, Ming Li

Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spok…