34 citations · 73 across the 14 of their papers we have counts for
12 papers · 1 filter
ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning
Sangho Lee, Jiwan Chung, Youngjae Yu +4
The natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the eve…
Learning Object Detection from Captions via Textual Scene Attributes
Achiya Jerbi, Roei Herzig, Jonathan Berant +2
Object detection is a fundamental task in computer vision, requiring large annotated datasets that are difficult to collect, as annotators need to label objects and their bounding…
A causal view of compositional zero-shot recognition
Yuval Atzmon, Felix Kreuk, Uri Shalit +1
People easily recognize new visual categories that are new combinations of known components. This compositional generalization capacity is critical for learning in real-world domai…
Contrastive Learning for Weakly Supervised Phrase Grounding
Tanmay Gupta, Arash Vahdat, Gal Chechik +3
Phrase grounding, the problem of associating image regions to caption words, is a crucial component of vision-language tasks. We show that phrase grounding can be learned by optimi…
Learning Object Permanence from Video
Aviv Shamsian, Ofri Kleinfeld, Amir Globerson +1
Object Permanence allows people to reason about the location of non-visible objects, by understanding that they continue to exist even when not perceived directly. Object Permanenc…
Learning Canonical Representations for Scene Graph to Image Generation
Roei Herzig, Amir Bar, Huijuan Xu +3
Generating realistic images of complex visual scenes becomes challenging when one wishes to control the structure of the generated images. Previous approaches showed that scenes wi…