16 citations · 32 across the 22 of their papers we have counts for
6 papers · 1 filter
Learning in your voice: Non-parallel voice conversion based on speaker consistency loss
Yoohwan Kwon, Soo-Whan Chung, Hee-Soo Heo +1
In this paper, we propose a novel voice conversion strategy to resolve the mismatch between the training and conversion scenarios when parallel speech corpus is unavailable for tra…
MIRNet: Learning multiple identities representations in overlapped speech
Hyewon Han, Soo-Whan Chung, Hong-Goo Kang
Many approaches can derive information about a single speaker's identity from the speech by learning to recognize consistent characteristics of acoustic parameters. However, it is…
Intra-class variation reduction of speaker representation in disentanglement framework
Yoohwan Kwon, Soo-Whan Chung, Hong-Goo Kang
In this paper, we propose an effective training strategy to ex-tract robust speaker representations from a speech signal. Oneof the key challenges in speaker recognition tasks is t…
FaceFilter: Audio-visual speech separation using still images
Soo-Whan Chung, Soyeon Choe, Joon Son Chung +1
The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that…
Seeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision
Soo-Whan Chung, Hong Goo Kang, Joon Son Chung
The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effec…
Improving LPCNet-based Text-to-Speech with Linear Prediction-structured Mixture Density Network
Min-Jae Hwang, Eunwoo Song, Ryuichi Yamamoto +2
In this paper, we propose an improved LPCNet vocoder using a linear prediction (LP)-structured mixture density network (MDN). The recently proposed LPCNet vocoder has successfully…