4 papers · 1 filter
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
Zhikang Niu, Shujie Hu, Jeongsoo Choi +8
Mel-spectrograms have been widely used in zero-shot text-to-speech (TTS); their inherent redundancy leads to inefficiency in text-speech alignment. Compact VAE-based latent represe…
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
Haitao Li, Chunxiang Jin, Chenglin Li +3
Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result,…
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
Mingyu Cui, Mengzhe Geng, Jiajun Deng +8
This paper investigates four types of cross-utterance speech contexts modeling approaches for streaming and non-streaming Conformer-Transformer (C-T) ASR systems: i) input audio fe…
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
Jeongsoo Choi, Zhikang Niu, Ji-Hoon Kim +3
The goal of this paper is to optimize the training process of diffusion-based text-to-speech models. While recent studies have achieved remarkable advancements, their training dema…