activity
20152022
most citedXCiT: Cross-Covariance Image Transformers

234 citations · 712 across the 22 of their papers we have counts for

collaborators
Showing cs.CVShow all

38 papers · 1 filter

cs.CV20224 cited

Multi-Task Learning of Object State Changes from Uncurated Videos

Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2

We aim to learn to temporally localize object state changes and the corresponding state-modifying actions by observing people interacting with objects in long uncurated web videos.…

cs.CV202214 cited

Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2

Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar ob…

cs.CV202212 cited

Weakly-supervised segmentation of referring expressions

Robin Strudel, Ivan Laptev, Cordelia Schmid

Visual grounding localizes regions (boxes or segments) in the image corresponding to given referring expressions. In this work we address image segmentation from referring expressi…

cs.CV20221 cited

Learning to Answer Visual Questions from Web Videos

Antoine Yang, Antoine Miech, Josef Sivic +2

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and preve…

cs.CV20222 cited

Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web Videos

Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2

Human actions often induce changes of object states such as "cutting an apple", "cleaning shoes" or "pouring coffee". In this paper, we seek to temporally localize object states (e…

cs.CV2022

Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation

Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2

Following language instructions to navigate in unseen environments is a challenging problem for autonomous embodied agents. The agent not only needs to ground languages in visual s…