3 papers
cs.SD2025
AdaptVC: High Quality Voice Conversion with Adaptive Learning
Jaehun Kim, Ji-Hoon Kim, Yeunju Choi +3
The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to e…
eess.AS2024
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
Jaemin Jung, Junseok Ahn, Chaeyoung Jung +3
We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for in…
eess.AS2024
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
Tan Dat Nguyen, Ji-Hoon Kim, Youngjoon Jang +2
The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key componen…