activity
20242026
most citedHierarchical Emotion Prediction and Control in Text-to-Speech Synthesis

16 citations · 27 across the 21 of their papers we have counts for

collaborators
Showing cs.SDShow all

12 papers · 1 filter

cs.SD2026

TokAN: Accent Normalization Using Self-Supervised Speech Tokens

Qibing Bai, Shuai Wang, Yuhan Du +3

Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity. The current techniques either require natura…

cs.SD2026

Borderless Long Speech Synthesis

Xingchen Song, Di Wu, Dinghao Zhou +12

Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both a…

cs.SD2026

AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow

Duojia Li, Shuhan Zhang, Zihan Qian +5

In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flo…

cs.SD2025

Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation

Wupeng Wang, Zexu Pan, Xinke Li +2

Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a pro…

cs.SD2025

Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation

Wupeng Wang, Zexu Pan, Jingru Lin +2

Speech separation seeks to isolate individual speech signals from a multi-talk speech mixture. Despite much progress, a system well-trained on synthetic data often experiences perf…

cs.SD2025

ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification

Yi Ma, Shuai Wang, Tianchi Liu +1

In speaker verification, we use computational method to verify if an utterance matches the identity of an enrolled speaker. This task is similar to the manual task of forensic voic…