12 papers
On the Usefulness of Diffusion-Based Room Impulse Response Interpolation to Microphone Array Processing
Sagi Della Torre, Mirco Pezzoli, Fabio Antonacci +1
Room Impulse Responses estimation is a fundamental problem in spatial audio processing and speech enhancement. In this paper, we build upon our previously introduced diffusion-base…
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
Francisco Messina, Francesca Ronchini, Luca Comanducci +2
A persistent challenge in generative audio models is data replication, where the model unintentionally generates parts of its training data during inference. In this work, we addre…
PAGURI: a user experience study of creative interaction with text-to-music models
Francesca Ronchini, Luca Comanducci, Gabriele Perego +1
In recent years, text-to-music models have been the biggest breakthrough in automatic music generation. While they are unquestionably a showcase of technological progress, it is no…
Training-Free Multimodal Guidance for Video to Audio Generation
Eleonora Grassucci, Giuliano Galadini, Giordano Cicchetti +3
Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, an…
AI-Assisted Music Production: A User Study on Text-to-Music Models
Francesca Ronchini, Luca Comanducci, Simone Marcucci +1
Text-to-music models have revolutionized the creative landscape, offering new possibilities for music creation. Yet their integration into musicians workflows remains underexplored…
Dynamic Real-Time Ambisonics Order Adaptation for Immersive Networked Music Performances
Paolo Ostan, Carlo Centofanti, Mirco Pezzoli +3
Advanced remote applications such as Networked Music Performance (NMP) require solutions to guarantee immersive real-world-like interaction among users. Therefore, the adoption of…