4 papers
Attention Alignment Between Humans and Vision-Language Models
Isaac R. Christian, Udith Haputhanthrige, Hanna Hornfeld +4
Visual perception depends on top-down goals and bottom-up sensory mechanisms. Vision-language models implement both, allowing us to treat each component as a separable hypothesis a…
Binding Visual Features Point by Point
Udith Haputhanthri, Declan Campbell, Rim Assouel +2
Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, including many tasks that are relat…
Few-Shot Learning of Visual Compositional Concepts through Probabilistic Schema Induction
Andrew Jun Lee, Taylor Webb, Trevor Bihl +2
The ability to learn new visual concepts from limited examples is a hallmark of human cognition. While traditional category learning models represent each example as an unstructure…
Evaluating Compositional Scene Understanding in Multimodal Generative Models
Shuhao Fu, Andrew Jun Lee, Anna Wang +4
The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential for computer vision systems to…