5 papers
PHALAR: Phasors for Learned Musical Audio Representations
Davide Marincione, Michele Mancusi, Giorgio Strano +4
Stem retrieval, the task of matching missing stems to a given audio submix, is a key challenge currently limited by models that discard temporal information. We introduce PHALAR, a…
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
Xinlei Niu, Kin Wai Cheuk, Jing Zhang +8
Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing me…
ITO-Master: Inference-Time Optimization for Audio Effects Modeling of Music Mastering Processors
Junghyun Koo, Marco A. MartÃnez-RamÃrez, Wei-Hsiang Liao +3
Music mastering style transfer aims to model and apply the mastering characteristics of a reference track to a target track, simulating the professional mastering process. However,…
High-Resolution Speech Restoration with Latent Diffusion Model
Tushar Dhyani, Florian Lux, Michele Mancusi +3
Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions fre…
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
Michele Mancusi, Yurii Halychanskyi, Kin Wai Cheuk +8
Music timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose…