8 citations · 8 across the 11 of their papers we have counts for
4 papers · 1 filter
Human-Like Coarse Object Representations in Vision Models
Andrey Gizdov, Andrea Procopio, Yichen Li +2
Humans appear to represent objects for intuitive physics with coarse, volumetric bodies'' that smooth concavities - trading fine visual details for efficient physical predictions -…
Towards aligned body representations in vision models
Andrey Gizdov, Andrea Procopio, Yichen Li +2
Human physical reasoning relies on internal "body" representations - coarse, volumetric approximations that capture an object's extent and support intuitive predictions about motio…
Chain of Time: In-Context Physical Simulation with Image Generation Models
YingQiao Wang, Eric Bigelow, Boyi Li +1
We propose a novel cognitively-inspired method to improve and interpret physical simulation in vision-language models. Our ``Chain of Time" method involves generating a series of i…
Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models
Colin Conwell, Rupert Tawiah-Quashie, Tomer Ullman
Despite remarkable progress in multi-modal AI research, there is a salient domain in which modern AI continues to lag considerably behind even human children: the reliable deployme…