158 citations · 242 across the 13 of their papers we have counts for
12 papers · 1 filter
Preference Optimization for Non-Verbal Vocalization Synthesis
Haoyang Li, Chenglin Xu, Junchuan Zhao +4
Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains po…
EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis
Tianchi Liu, Zeyang Song, Tianrui Wang +3
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotion…
L-SpEx: Localized Target Speaker Extraction
Meng Ge, Chenglin Xu, Longbiao Wang +3
Speaker extraction aims to extract the target speaker's voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction…
Target Speaker Verification with Selective Auditory Attention for Single and Multi-talker Speech
Chenglin Xu, Wei Rao, Jibin Wu +1
Speaker verification has been studied mostly under the single-talker condition. It is adversely affected in the presence of interference speakers. Inspired by the study on target s…
Multi-stage Speaker Extraction with Utterance and Frame-Level Reference Signals
Meng Ge, Chenglin Xu, Longbiao Wang +3
Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extract…
Muse: Multi-modal target speaker extraction with visual cues
Zexu Pan, Ruijie Tao, Chenglin Xu +1
Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. O…