1.3k citations · 1.9k across the 19 of their papers we have counts for
6 papers · 2 filters
Multi-Task Learning of Object State Changes from Uncurated Videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2
We aim to learn to temporally localize object state changes and the corresponding state-modifying actions by observing people interacting with objects in long uncurated web videos.…
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models
Antoine Yang, Antoine Miech, Josef Sivic +2
Video question answering (VideoQA) is a complex task that requires diverse multi-modal data for training. Manual annotation of question and answers for videos, however, is tedious…
Learning to Answer Visual Questions from Web Videos
Antoine Yang, Antoine Miech, Josef Sivic +2
Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and preve…
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc +24
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Fl…
Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web Videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2
Human actions often induce changes of object states such as "cutting an apple", "cleaning shoes" or "pouring coffee". In this paper, we seek to temporally localize object states (e…
TubeDETR: Spatio-Temporal Video Grounding with Transformers
Antoine Yang, Antoine Miech, Josef Sivic +2
We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeli…