works on

From the 2 of 17 linked papers with an AI index.

activity
20242026
collaborators

17 papers

eess.AS2026

WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction

Ke Zhang, Xiaoyang Yu, Haoyu Li +3

WeSep is a modular framework that treats target speaker extraction as a cue‑conditioned learning problem, separating cue modules from the separator backbone to flexibly incorporate…

eess.AS2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Shuai Wang, Zihan Qian, Ke Zhang +9

The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…

eess.AS2026

Detect, Attend and Extract: Keyword Guided Target Speaker Extraction

Haoyu Li, Yu Xi, Yidi Jiang +5

Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on…

cs.SD2026

AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow

Duojia Li, Shuhan Zhang, Zihan Qian +5

In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flo…

eess.AS2025

SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement

Chenyu Yang, Shuai Wang, Hangting Chen +3

Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-base…

eess.AS2025

Direct Preference Optimization for Speech Autoregressive Diffusion Models

Zhijun Liu, Dongya Jia, Xiaoqiang Wang +4

Autoregressive diffusion models (ARDMs) have recently been applied to speech generation, achieving state-of-the-art (SOTA) performance in zero-shot text-to-speech. By autoregressiv…