6 papers
Generative Models for Helmholtz Equation Solutions: A Dataset of Acoustic Materials
Riccardo Fosco Gramaccioni, Christian Marinoni, Fabrizio Frezza +2
Accurate simulation of wave propagation in complex acoustic materials is crucial for applications in sound design, noise control, and material engineering. Traditional numerical so…
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
Christian Marinoni, Riccardo Fosco Gramaccioni, Eleonora Grassucci +1
The generation of sounding videos has seen significant advancements with the advent of diffusion models. However, existing methods often lack the fine-grained control needed to gen…
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
Riccardo Fosco Gramaccioni, Christian Marinoni, Eleonora Grassucci +3
In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on…
StereoSync: Spatially-Aware Stereo Audio Generation from Video
Christian Marinoni, Riccardo Fosco Gramaccioni, Kazuki Shimada +3
Although audio generation has been widely studied over recent years, video-aligned audio generation still remains a relatively unexplored frontier. To address this gap, we introduc…
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Riccardo Fosco Gramaccioni, Christian Marinoni, Emilian Postolache +4
Traditional sound design workflows rely on manual alignment of audio events to visual cues, as in Foley sound design, where everyday actions like footsteps or object interactions a…
Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality
Christian Marinoni, Riccardo Fosco Gramaccioni, Changan Chen +2
The primary goal of the L3DAS23 Signal Processing Grand Challenge at ICASSP 2023 is to promote and support collaborative research on machine learning for 3D audio signal processing…