4 papers
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Riccardo Fosco Gramaccioni, Christian Marinoni, Emilian Postolache +4
Traditional sound design workflows rely on manual alignment of audio events to visual cues, as in Foley sound design, where everyday actions like footsteps or object interactions a…
Naturalistic Music Decoding from EEG Data via Latent Diffusion Models
Emilian Postolache, Natalia Polouliakh, Hiroaki Kitano +4
In this article, we explore the potential of using latent diffusion models, a family of powerful generative models, for the task of reconstructing naturalistic music from electroen…
COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
Ruben Ciranni, Giorgio Mariani, Michele Mancusi +4
We present COCOLA (Coherence-Oriented Contrastive Learning for Audio), a contrastive learning method for musical audio representations that captures the harmonic and rhythmic coher…
Zero-Shot Duet Singing Voices Separation with Diffusion Models
Chin-Yun Yu, Emilian Postolache, Emanuele Rodolà +1
In recent studies, diffusion models have shown promise as priors for solving audio inverse problems. These models allow us to sample from the posterior distribution of a target sig…