8 papers
Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
Sounak Mondal, Dimitris Samaras, Gregory Zelinsky +1
We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as…
Pathologist Attention-Aligned Report Generation for Prostate Histopathology
Ruoyu Xue, Suryakant Singh, Souradeep Chakraborty +12
The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracted from whole-slide images (W…
Human-like Object Grouping in Self-supervised Vision Transformers
Hossein Adeli, Seoyoung Ahn, Andrew Luo +3
Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their…
LVLMs and Humans Ground Differently in Referential Communication
Peter Zeng, Weiling Li, Amie J. Paige +6
For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical. But this ability to collaborate remains limited by a c…
Generating metamers of human scene understanding
Ritik Raina, Abe Leite, Alexandros Graikos +3
Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent underst…
Personalized Image Descriptions from Attention Sequences
Ruoyu Xue, Hieu Le, Jingyi Xu +5
People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to s…