4 papers
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
Minhyeok Yun, Yong-Hoon Choi
Singing voice synthesis (SVS) aims to generate natural and expressive singing waveforms from symbolic musical scores. In cVAE-based SVS, however, a mismatch arises because the deco…
Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
Euiyeon Kim, Yong-Hoon Choi
We introduce a new music source separation model tailored for accurate vocal isolation. Unlike Transformer-based approaches, which often fail to capture intermittently occurring vo…
Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing
Sukhyun Jeong, Yong-Hoon Choi
Text-based 3D motion generation aims to automatically synthesize diverse motions from natural-language descriptions to extend user creativity, whereas motion editing modifies an ex…
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
Seongho Hong, Yong-Hoon Choi
While transformers demonstrate outstanding performance across various audio tasks, their application to neural vocoders remains challenging. Neural vocoders require the generation…