1 paper
Santiago Pascual, Chunghsin Yeh, Ioannis Tsiamas +1
Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visua…