4 papers
Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
Changjin Han, Seokgi Lee, Gyuhyeon Nam +1
Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and soun…
KoDF: A Large-scale Korean DeepFake Detection Dataset
Patrick Kwon, Jaeseong You, Gyuhyeon Nam +2
A variety of effective face-swap and face-reenactment methods have been publicized in recent years, democratizing the face synthesis technology to a great extent. Videos generated…
GAN Vocoder: Multi-Resolution Discriminator Is All You Need
Jaeseong You, Dalhyun Kim, Gyuhyeon Nam +2
Several of the latest GAN-based vocoders show remarkable achievements, outperforming autoregressive and flow-based competitors in both qualitative and quantitative measures while s…
Axial Residual Networks for CycleGAN-based Voice Conversion
Jaeseong You, Gyuhyeon Nam, Dalhyun Kim +1
We propose a novel architecture and improved training objectives for non-parallel voice conversion. Our proposed CycleGAN-based model performs a shape-preserving transformation dir…