65 citations · 142 across the 10 of their papers we have counts for
4 papers · 1 filter
FoodSense: A Multisensory Food Dataset and Benchmark for Predicting Taste, Smell, Texture, and Sound from Images
Sabab Ishraq, Aarushi Aarushi, Juncai Jiang +1
Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has fo…
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
Hritik Bansal, Po-Nien Kung, P. Jeffrey Brantingham +2
Multimodal event argument role labeling (EARL), a task that assigns a role for each event participant (object) in an image is a complex challenge. It requires reasoning over the en…
Compressed Vision for Efficient Video Understanding
Olivia Wiles, Joao Carreira, Iain Barr +2
Experience and reasoning occur across multiple temporal scales: milliseconds, seconds, hours or days. The vast majority of computer vision research, however, still focuses on indiv…
Transframer: Arbitrary Frame Prediction with Generative Models
Charlie Nash, João Carreira, Jacob Walker +4
We present a general-purpose framework for image modelling and vision tasks based on probabilistic frame prediction. Our approach unifies a broad range of tasks, from image segment…