6 papers
Understanding Pan-Sharpening via Generalized Inverse
Shiqi Liu, Yihua Tan, Yutong Bai +1
Pan-sharpening algorithms utilize a panchromatic image and a multispectral image to generate a high spatial and high spectral image. However, the optimizations of the algorithms ar…
KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
Eunice Yiu, Maan Qraitem, Anisa Noor Majhi +5
This paper investigates visual analogical reasoning in large multimodal models (LMMs) compared to human adults and children. A "visual analogy" is an abstract rule inferred from on…
Analyzing The Language of Visual Tokens
David M. Chan, Rodolfo Corona, Joonyong Park +3
With the introduction of transformer-based models for vision and language tasks, such as LLaVA and Chameleon, there has been renewed interest in the discrete tokenized representati…
Finding Visual Task Vectors
Alberto Hojel, Yutong Bai, Trevor Darrell +2
Visual Prompting is a technique for teaching models to perform a visual task via in-context examples, without any additional training. In this work, we analyze the activations of M…
Evaluating Multiview Object Consistency in Humans and Image Models
Tyler Bonnen, Stephanie Fu, Yutong Bai +5
We introduce a benchmark to directly evaluate the alignment between human observers and vision models on a 3D shape inference task. We leverage an experimental design from the cogn…
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
Dantong Niu, Yuvan Sharma, Giscard Biamby +5
In recent years, instruction-tuned Large Multimodal Models (LMMs) have been successful at several tasks, including image captioning and visual question answering; yet leveraging th…