activity
20172026
most citedThinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers

133 citations · 229 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

FoodSense: A Multisensory Food Dataset and Benchmark for Predicting Taste, Smell, Texture, and Sound from Images

Sabab Ishraq, Aarushi Aarushi, Juncai Jiang +1

Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has fo…

cs.CV20224 cited

Multi-Task Learning of Object State Changes from Uncurated Videos

Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2

We aim to learn to temporally localize object state changes and the corresponding state-modifying actions by observing people interacting with objects in long uncurated web videos.…

cs.CV20221 cited

Learning to Answer Visual Questions from Web Videos

Antoine Yang, Antoine Miech, Josef Sivic +2

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and preve…

cs.CV20222 cited

Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web Videos

Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2

Human actions often induce changes of object states such as "cutting an apple", "cleaning shoes" or "pouring coffee". In this paper, we seek to temporally localize object states (e…

cs.CV2021133 cited

Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers

Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev +2

Our objective is language-based search of large-scale image and video datasets. For this task, the approach that consists of independently mapping text and vision to a joint embedd…

cs.CV2020

Just Ask: Learning to Answer Questions from Millions of Narrated Videos

Antoine Yang, Antoine Miech, Josef Sivic +2

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and preve…