16 citations · 27 across the 21 of their papers we have counts for
12 papers · 1 filter
TokAN: Accent Normalization Using Self-Supervised Speech Tokens
Qibing Bai, Shuai Wang, Yuhan Du +3
Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity. The current techniques either require natura…
Borderless Long Speech Synthesis
Xingchen Song, Di Wu, Dinghao Zhou +12
Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both a…
AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow
Duojia Li, Shuhan Zhang, Zihan Qian +5
In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flo…
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
Wupeng Wang, Zexu Pan, Xinke Li +2
Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a pro…
Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation
Wupeng Wang, Zexu Pan, Jingru Lin +2
Speech separation seeks to isolate individual speech signals from a multi-talk speech mixture. Despite much progress, a system well-trained on synthetic data often experiences perf…
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
Yi Ma, Shuai Wang, Tianchi Liu +1
In speaker verification, we use computational method to verify if an utterance matches the identity of an enrolled speaker. This task is similar to the manual task of forensic voic…