180 citations
- NAVER Cloud (South Korea)KR26 papers
- Seoul National UniversityKR8 papers
- Yonsei UniversityKR8 papers
- Sungkyunkwan UniversityKR6 papers
- Korea Advanced Institute of Science and TechnologyKR5 papers
- Virginia TechUS5 papers
- Line Corporation (Japan)JP4 papers
- Korea UniversityKR3 papers
- Daegu Gyeongbuk Institute of Science and TechnologyKR2 papers
- Inha UniversityKR2 papers
- Pusan National UniversityKR2 papers
- Space Solutions (South Korea)KR2 papers
7 papers · 1 filter
An empirical study on speech restoration guided by self supervised speech representation
Jaeuk Byun, Youna Ji, Soo Whan Chung +2
Enhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping…
Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation
Hemlata Tak, Massimiliano Todisco, Xin Wang +3
The performance of spoofing countermeasure systems depends fundamentally upon the use of sufficiently representative training data. With this usually being limited, current solutio…
Three-class Overlapped Speech Detection using a Convolutional Recurrent Neural Network
Jee-weon Jung, Hee-Soo Heo, Youngki Kwon +2
In this work, we propose an overlapped speech detection system trained as a three-class classifier. Unlike conventional systems that perform binary classification as to whether or…
Improved parallel WaveGAN vocoder with perceptually weighted spectrogram loss
Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang +3
This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder success…
TTS-by-TTS: TTS-driven Data Augmentation for Fast and High-Quality Speech Synthesis
Min-Jae Hwang, Ryuichi Yamamoto, Eunwoo Song +1
In this paper, we propose a text-to-speech (TTS)-driven data augmentation method for improving the quality of a non-autoregressive (AR) TTS system. Recently proposed non-AR models,…
Improving LPCNet-based Text-to-Speech with Linear Prediction-structured Mixture Density Network
Min-Jae Hwang, Eunwoo Song, Ryuichi Yamamoto +2
In this paper, we propose an improved LPCNet vocoder using a linear prediction (LP)-structured mixture density network (MDN). The recently proposed LPCNet vocoder has successfully…