2 papers
eess.AS2022
Adversarial Speaker-Consistency Learning Using Untranscribed Speech Data for Zero-Shot Multi-Speaker Text-to-Speech
Byoung Jin Choi, Myeonghun Jeong, Minchan Kim +2
Several recently proposed text-to-speech (TTS) models achieved to generate the speech samples with the human-level quality in the single-speaker and multi-speaker TTS scenarios wit…
eess.AS2022
Fully Unsupervised Training of Few-shot Keyword Spotting
Dongjune Lee, Minchan Kim, Sung Hwan Mun +2
For training a few-shot keyword spotting (FS-KWS) model, a large labeled dataset containing massive target keywords has known to be essential to generalize to arbitrary target keyw…