93 citations · 225 across the 20 of their papers we have counts for
5 papers · 1 filter
Analyzing The Language of Visual Tokens
David M. Chan, Rodolfo Corona, Joonyong Park +3
With the introduction of transformer-based models for vision and language tasks, such as LLaVA and Chameleon, there has been renewed interest in the discrete tokenized representati…
Evaluating Multiview Object Consistency in Humans and Image Models
Tyler Bonnen, Stephanie Fu, Yutong Bai +5
We introduce a benchmark to directly evaluate the alignment between human observers and vision models on a 3D shape inference task. We leverage an experimental design from the cogn…
KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
Eunice Yiu, Maan Qraitem, Anisa Noor Majhi +5
This paper investigates visual analogical reasoning in large multimodal models (LMMs) compared to human adults and children. A "visual analogy" is an abstract rule inferred from on…
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
Dantong Niu, Yuvan Sharma, Giscard Biamby +5
In recent years, instruction-tuned Large Multimodal Models (LMMs) have been successful at several tasks, including image captioning and visual question answering; yet leveraging th…
Finding Visual Task Vectors
Alberto Hojel, Yutong Bai, Trevor Darrell +2
Visual Prompting is a technique for teaching models to perform a visual task via in-context examples, without any additional training. In this work, we analyze the activations of M…