1 citations · 1 across the 4 of their papers we have counts for
4 papers · 1 filter
MCR-Data2vec 2.0: Improving Self-supervised Speech Pre-training via Model-level Consistency Regularization
Ji Won Yoon, Seok Min Kim, Nam Soo Kim
Self-supervised learning (SSL) has shown significant progress in speech processing tasks. However, despite the intrinsic randomness in the Transformer structure, such as dropout va…
Inter-KD: Intermediate Knowledge Distillation for CTC-Based Automatic Speech Recognition
Ji Won Yoon, Beom Jun Woo, Sunghwan Ahn +2
Recently, the advance in deep learning has brought a considerable improvement in the end-to-end speech recognition field, simplifying the traditional pipeline while producing promi…
Adversarial Speaker-Consistency Learning Using Untranscribed Speech Data for Zero-Shot Multi-Speaker Text-to-Speech
Byoung Jin Choi, Myeonghun Jeong, Minchan Kim +2
Several recently proposed text-to-speech (TTS) models achieved to generate the speech samples with the human-level quality in the single-speaker and multi-speaker TTS scenarios wit…
Fully Unsupervised Training of Few-shot Keyword Spotting
Dongjune Lee, Minchan Kim, Sung Hwan Mun +2
For training a few-shot keyword spotting (FS-KWS) model, a large labeled dataset containing massive target keywords has known to be essential to generalize to arbitrary target keyw…