5 papers · 1 filter
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Riccardo Fosco Gramaccioni, Christian Marinoni, Emilian Postolache +4
Traditional sound design workflows rely on manual alignment of audio events to visual cues, as in Foley sound design, where everyday actions like footsteps or object interactions a…
STAGE: Stemmed Accompaniment Generation through Prefix-Based Conditioning
Giorgio Strano, Chiara Ballanti, Donato Crisostomi +3
Recent advances in generative models have made it possible to create high-quality, coherent music, with some systems delivering production-level output. Yet, most existing models f…
COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
Ruben Ciranni, Giorgio Mariani, Michele Mancusi +4
We present COCOLA (Coherence-Oriented Contrastive Learning for Audio), a contrastive learning method for musical audio representations that captures the harmonic and rhythmic coher…
Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models
Emilian Postolache, Giorgio Mariani, Luca Cosmo +2
Multi-Source Diffusion Models (MSDM) allow for compositional musical generation tasks: generating a set of coherent sources, creating accompaniments, and performing source separati…
Multi-Source Diffusion Models for Simultaneous Music Generation and Separation
Giorgio Mariani, Irene Tallini, Emilian Postolache +3
In this work, we define a diffusion-based generative model capable of both music synthesis and source separation by learning the score of the joint probability density of sources s…