2 citations · 2 across the 3 of their papers we have counts for
3 papers · 1 filter
StereoFoley: Object-Aware Stereo Audio Generation from Video
Tornike Karchkhadze, Kuan-Lin Chen, Mojtaba Heydari +4
We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While rece…
ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model
Mojtaba Heydari, Mehrez Souden, Bruno Conejo +1
We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sou…
Resource-constrained stereo singing voice cancellation
Clara Borrelli, James Rae, Dogac Basaran +3
We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore…