6 papers
VocalCoachBench: Benchmarking Audio-Language Models on Expert Feedback for Singing
Hayeon Bang, Hounsu Kim, Wonil Kim +1
Recent audio-language models are increasingly evaluated on recognizing, describing, and reasoning about audio, but expert-facing applications require a different capability: produc…
SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion
Hounsu Kim, Juhan Nam
Speaker-decoupled speech codecs can reduce bitrate by separating global speaker attributes from local content and prosody, while supporting voice conversion. Existing speaker-decou…
D3PIA: A Discrete Denoising Diffusion Model for Piano Accompaniment Generation From Lead sheet
Eunjin Choi, Hounsu Kim, Hayeon Bang +2
Generating piano accompaniments in the symbolic music domain is a challenging task that requires producing a complete piece of piano music from given melody and chord constraints,…
D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription
Hounsu Kim, Taegyun Kwon, Juhan Nam
Diffusion models have been widely used in the generative domain due to their convincing performance in modeling complex data distributions. Moreover, they have shown competitive re…
CONMOD: Controllable Neural Frame-based Modulation Effects
Gyubin Lee, Hounsu Kim, Junwon Lee +1
Deep learning models have seen widespread use in modelling LFO-driven audio effects, such as phaser and flanger. Although existing neural architectures exhibit high-quality emulati…
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
Hounsu Kim, Soonbeom Choi, Juhan Nam
Synthesizing performing guitar sound is a highly challenging task due to the polyphony and high variability in expression. Recently, deep generative models have shown promising res…