From the 2 of 17 linked papers with an AI index.
17 papers
WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction
Ke Zhang, Xiaoyang Yu, Haoyu Li +3
WeSep is a modular framework that treats target speaker extraction as a cue‑conditioned learning problem, separating cue modules from the separator backbone to flexibly incorporate…
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Shuai Wang, Zihan Qian, Ke Zhang +9
The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…
Detect, Attend and Extract: Keyword Guided Target Speaker Extraction
Haoyu Li, Yu Xi, Yidi Jiang +5
Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on…
AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow
Duojia Li, Shuhan Zhang, Zihan Qian +5
In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flo…
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
Chenyu Yang, Shuai Wang, Hangting Chen +3
Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-base…
Direct Preference Optimization for Speech Autoregressive Diffusion Models
Zhijun Liu, Dongya Jia, Xiaoqiang Wang +4
Autoregressive diffusion models (ARDMs) have recently been applied to speech generation, achieving state-of-the-art (SOTA) performance in zero-shot text-to-speech. By autoregressiv…