211 citations · 552 across the 47 of their papers we have counts for
92 papers · 1 filter
Learning text-to-video retrieval from image captioning
Lucas Ventura, Cordelia Schmid, Gül Varol
We describe a protocol to study text-to-video retrieval training with unlabeled videos, where we assume (i) no access to labels for any videos, i.e., no access to the set of ground…
POCO: 3D Pose and Shape Estimation with Confidence
Sai Kumar Dwivedi, Cordelia Schmid, Hongwei Yi +2
The regression of 3D Human Pose and Shape (HPS) from an image is becoming increasingly accurate. This makes the results useful for downstream tasks like human action recognition or…
UnLoc: A Unified Framework for Video Localization Tasks
Shen Yan, Xuehan Xiong, Arsha Nagrani +5
While large-scale image-text pretrained models such as CLIP have been used for multiple video-level tasks on trimmed videos, their use for temporal localization in untrimmed videos…
Object Goal Navigation with Recursive Implicit Maps
Shizhe Chen, Thomas Chabal, Ivan Laptev +1
Object goal navigation aims to navigate an agent to locations of a given object category in unseen environments. Classical methods explicitly build maps of environments and require…
CoVR-2: Automatic Data Construction for Composed Video Retrieval
Lucas Ventura, Antoine Yang, Cordelia Schmid +1
Composed Image Retrieval (CoIR) has recently gained popularity as a task that considers both text and image queries together, to search for relevant images in a database. Most CoIR…
Does Visual Pretraining Help End-to-End Reasoning?
Chen Sun, Calvin Luo, Xingyi Zhou +2
We aim to investigate whether end-to-end learning of visual reasoning can be achieved with general-purpose neural networks, with the help of visual pretraining. A positive result w…