3 papers
cs.LG2025
LauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language Models
Beilong Tang, Bang Zeng, Ming Li
We propose LauraTSE, an Auto-Regressive Decoder-Only Language Model for Target Speaker Extraction built upon the LauraGPT backbone. LauraTSE employs a small-scale auto-regressive d…
eess.AS2025
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
Bang Zeng, Ming Li
Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spok…
cs.SD2024
TSELM: Target Speaker Extraction using Discrete Tokens and Language Models
Beilong Tang, Bang Zeng, Ming Li
We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input token…