activity
20172026
most citedFlamingo: a Visual Language Model for Few-Shot Learning

1.3k citations · 1.9k across the 19 of their papers we have counts for

collaborators
Showing 2022 · cs.CVShow all

6 papers · 2 filters

cs.CV2022★ 4 cited

Multi-Task Learning of Object State Changes from Uncurated Videos

Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2

We aim to learn to temporally localize object state changes and the corresponding state-modifying actions by observing people interacting with objects in long uncurated web videos.…

cs.CV2022★ 64 cited

Zero-Shot Video Question Answering via Frozen Bidirectional Language Models

Antoine Yang, Antoine Miech, Josef Sivic +2

Video question answering (VideoQA) is a complex task that requires diverse multi-modal data for training. Manual annotation of question and answers for videos, however, is tedious…

cs.CV2022★ 1 cited

Learning to Answer Visual Questions from Web Videos

Antoine Yang, Antoine Miech, Josef Sivic +2

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and preve…

cs.CV2022★ 1.3k cited

Flamingo: a Visual Language Model for Few-Shot Learning

Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc +24

Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Fl…

cs.CV2022★ 2 cited

Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web Videos

Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2

Human actions often induce changes of object states such as "cutting an apple", "cleaning shoes" or "pouring coffee". In this paper, we seek to temporally localize object states (e…

cs.CV2022

TubeDETR: Spatio-Temporal Video Grounding with Transformers

Antoine Yang, Antoine Miech, Josef Sivic +2

We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeli…