82 citations · 83 across the 4 of their papers we have counts for
4 papers
Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity
Santiago Pascual, Chunghsin Yeh, Ioannis Tsiamas +1
Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visua…
GASS: Generalizing Audio Source Separation with Large-scale Data
Jordi Pons, Xiaoyu Liu, Santiago Pascual +1
Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the pote…
CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models
Hao-Wen Dong, Xiaoyu Liu, Jordi Pons +5
Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acqu…
Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
Alberto Montes, Amaia Salvador, Santiago Pascual +1
This thesis explore different approaches using Convolutional and Recurrent Neural Networks to classify and temporally localize activities on videos, furthermore an implementation t…