133 citations · 229 across the 11 of their papers we have counts for
12 papers · 1 filter
FoodSense: A Multisensory Food Dataset and Benchmark for Predicting Taste, Smell, Texture, and Sound from Images
Sabab Ishraq, Aarushi Aarushi, Juncai Jiang +1
Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has fo…
Multi-Task Learning of Object State Changes from Uncurated Videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2
We aim to learn to temporally localize object state changes and the corresponding state-modifying actions by observing people interacting with objects in long uncurated web videos.…
Learning to Answer Visual Questions from Web Videos
Antoine Yang, Antoine Miech, Josef Sivic +2
Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and preve…
Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web Videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech +2
Human actions often induce changes of object states such as "cutting an apple", "cleaning shoes" or "pouring coffee". In this paper, we seek to temporally localize object states (e…
Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers
Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev +2
Our objective is language-based search of large-scale image and video datasets. For this task, the approach that consists of independently mapping text and vision to a joint embedd…
Just Ask: Learning to Answer Questions from Millions of Narrated Videos
Antoine Yang, Antoine Miech, Josef Sivic +2
Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and preve…