11 papers
Music Restoration via Latent Operator Optimization and Diffusion Model Priors
Michal Å vento, Eloi Moliner, Valtteri Kallinen +3
Music restoration seeks to recover a clean signal from an observed recording degraded by an unknown effect, distortion, or corruption. Existing systems often rely on paired trainin…
Frequency-Aware Self-Supervised Music Representation Learning
Yicheng Gu, Junan Zhang, Jerry Li +2
Self-supervised learning (SSL) has emerged as an essential paradigm for music information retrieval (MIR). While current SSL models achieve state-of-the-art performance across vari…
Aliasing-Free Neural Audio Synthesis
Yicheng Gu, Junan Zhang, Chaoren Wang +3
In neural audio synthesis, neural vocoders and codecs are models that reconstruct waveforms from acoustic and latent representations, which are essential to the resulting audio qua…
Deep Regularized RNNs for Virtual Analog Modeling
V. Valtteri Kallinen, Lauri Juvela, Thom Sherson
Virtual analog (VA) modeling methods seek to emulate analog audio hardware using digital signal processing (DSP). Modeling approaches fall into three broad categories: white-box me…
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
Yicheng Gu, Pablo Pérez Zarazaga, Chaoren Wang +4
Formant synthesis aims to generate speech with controllable formant structures, enabling precise control of vocal resonance and phonetic features. However, while existing formant s…
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
Zirui Li, Jens Edlund, Yicheng Gu +3
Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages. We present Nord-Pa…