4 papers
Human-like Object Grouping in Self-supervised Vision Transformers
Hossein Adeli, Seoyoung Ahn, Andrew Luo +3
Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their…
Generating metamers of human scene understanding
Ritik Raina, Abe Leite, Alexandros Graikos +3
Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent underst…
Look Hear: Gaze Prediction for Speech-directed Human Attention
Sounak Mondal, Seoyoung Ahn, Zhibo Yang +4
For computer systems to effectively interact with humans using spoken language, they need to understand how the words being generated affect the users' moment-by-moment attention.…
Predicting Visual Attention in Graphic Design Documents
Souradeep Chakraborty, Zijun Wei, Conor Kelton +4
We present a model for predicting visual attention during the free viewing of graphic design documents. While existing works on this topic have aimed at predicting static saliency…