7 citations · 10 across the 7 of their papers we have counts for
7 papers
HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis
Sang-Hoon Lee, Ha-Yeong Choi, Seung-Bin Kim +1
Large language models (LLM)-based speech synthesis has been widely adopted in zero-shot speech synthesis. However, they require a large-scale data and possess the same limitations…
PauseSpeech: Natural Speech Synthesis via Pre-trained Language Model and Pause-based Prosody Modeling
Ji-Sang Hwang, Sang-Hoon Lee, Seong-Whan Lee
Although text-to-speech (TTS) systems have significantly improved, most TTS systems still have limitations in synthesizing speech with appropriate phrasing. For natural speech synt…
GC-TTS: Few-shot Speaker Adaptation with Geometric Constraints
Ji-Hoon Kim, Sang-Hoon Lee, Ji-Hyun Lee +2
Few-shot speaker adaptation is a specific Text-to-Speech (TTS) system that aims to reproduce a novel speaker's voice with a few training data. While numerous attempts have been mad…
Fre-GAN: Adversarial Frequency-consistent Audio Synthesis
Ji-Hoon Kim, Sang-Hoon Lee, Ji-Hyun Lee +1
Although recent works on neural vocoder have improved the quality of synthesized audio, there still exists a gap between generated and ground-truth audio in frequency space. This d…
Reinforce-Aligner: Reinforcement Alignment Search for Robust End-to-End Text-to-Speech
Hyunseung Chung, Sang-Hoon Lee, Seong-Whan Lee
Text-to-speech (TTS) synthesis is the process of producing synthesized speech from text or phoneme input. Traditional TTS models contain multiple processing steps and require exter…
Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech Synthesis
Sang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh +2
While generative adversarial networks (GANs) based neural text-to-speech (TTS) systems have shown significant improvement in neural speech synthesis, there is no TTS system to lear…