4 papers
Large-scale Speaker Retrieval on Random Speaker Variability Subspace
Suwon Shon, Younggun Lee, Taesu Kim
This paper describes a fast speaker search system to retrieve segments of the same voice identity in the large-scale data. A recent study shows that Locality Sensitive Hashing (LSH…
Learning pronunciation from a foreign language in speech synthesis networks
Younggun Lee, Suwon Shon, Taesu Kim
Although there are more than 6,500 languages in the world, the pronunciations of many phonemes sound similar across the languages. When people learn a foreign language, their pronu…
Robust and fine-grained prosody control of end-to-end speech synthesis
Younggun Lee, Taesu Kim
We propose prosody embeddings for emotional and expressive speech synthesis networks. The proposed methods introduce temporal structures in the embedding networks, thus enabling fi…
Voice Imitating Text-to-Speech Neural Networks
Younggun Lee, Taesu Kim, Soo-Young Lee
We propose a neural text-to-speech (TTS) model that can imitate a new speaker's voice using only a small amount of speech sample. We demonstrate voice imitation using only a 6-seco…