4 papers
MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling
Yoonjeong Park, Jaekwon Im, Juhan Nam
Text-based singing voice editing (SVE) aims to revise sung lyrics while preserving the original melody, total duration, and non-edited regions. In this paper, we propose MeloDISing…
SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution
Jaekwon Im, Juhan Nam
Versatile audio super-resolution (SR) aims to predict high-frequency components from low-resolution audio across diverse domains such as speech, music, and sound effects. Existing…
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
Junwon Lee, Jaekwon Im, Dabin Kim +1
Foley sound synthesis is crucial for multimedia production, enhancing user experience by synchronizing audio and video both temporally and semantically. Recent studies on automatin…
FlashSR: One-step Versatile Audio Super-resolution via Diffusion Distillation
Jaekwon Im, Juhan Nam
Versatile audio super-resolution (SR) is the challenging task of restoring high-frequency components from low-resolution audio with sampling rates between 4kHz and 32kHz in various…