25 citations · 91 across the 22 of their papers we have counts for
7 papers · 1 filter
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
Haibin Wu, Xiaofei Wang, Sefik Emre Eskimez +8
People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) syste…
Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker +9
Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting t…
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Zeqian Ju, Yuancheng Wang, Kai Shen +16
While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intric…
t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability
Jian Wu, Naoyuki Kanda, Takuya Yoshioka +3
Token-level serialized output training (t-SOT) was recently proposed to address the challenge of streaming multi-talker automatic speech recognition (ASR). T-SOT effectively handle…
Pre-training End-to-end ASR Models with Augmented Speech Samples Queried by Text
Eric Sun, Jinyu Li, Jian Xue +1
In end-to-end automatic speech recognition system, one of the difficulties for language expansion is the limited paired speech and text training data. In this paper, we propose a n…
Speaker Change Detection for Transformer Transducer ASR
Jian Wu, Zhuo Chen, Min Hu +2
Speaker change detection (SCD) is an important feature that improves the readability of the recognized words from an automatic speech recognition (ASR) system by breaking the word…