1 citations · 1 across the 1 of their papers we have counts for
3 papers
Revealing Occlusions with 4D Neural Fields
Basile Van Hoorick, Purva Tendulkar, Didac Suris +3
For computer vision systems to operate in dynamic situations, they need to be able to represent and reason about object permanence. We introduce a framework for learning to estimat…
Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input
David Harwath, Adrià Recasens, Dídac Surís +3
In this paper, we explore neural network models that learn to associate segments of spoken audio captions with the semantically relevant portions of natural images that they refer…
Cross-modal Embeddings for Video and Audio Retrieval
Didac Surís, Amanda Duarte, Amaia Salvador +2
The increasing amount of online videos brings several opportunities for training self-supervised neural networks. The creation of large scale datasets of videos such as the YouTube…