234 citations · 712 across the 22 of their papers we have counts for
38 papers · 1 filter
Multi-Task Learning of Object State Changes from Uncurated Videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2
We aim to learn to temporally localize object state changes and the corresponding state-modifying actions by observing people interacting with objects in long uncurated web videos.…
Language Conditioned Spatial Relation Reasoning for 3D Object Grounding
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2
Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar ob…
Weakly-supervised segmentation of referring expressions
Robin Strudel, Ivan Laptev, Cordelia Schmid
Visual grounding localizes regions (boxes or segments) in the image corresponding to given referring expressions. In this work we address image segmentation from referring expressi…
Learning to Answer Visual Questions from Web Videos
Antoine Yang, Antoine Miech, Josef Sivic +2
Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and preve…
Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web Videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2
Human actions often induce changes of object states such as "cutting an apple", "cleaning shoes" or "pouring coffee". In this paper, we seek to temporally localize object states (e…
Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2
Following language instructions to navigate in unseen environments is a challenging problem for autonomous embodied agents. The agent not only needs to ground languages in visual s…