1 citations · 1 across the 3 of their papers we have counts for
3 papers · 1 filter
Attention Alignment Between Humans and Vision-Language Models
Isaac R. Christian, Udith Haputhanthrige, Hanna Hornfeld +4
Visual perception depends on top-down goals and bottom-up sensory mechanisms. Vision-language models implement both, allowing us to treat each component as a separable hypothesis a…
Binding Visual Features Point by Point
Udith Haputhanthri, Declan Campbell, Rim Assouel +2
Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, including many tasks that are relat…
Visual symbolic mechanisms: Emergent symbol processing in vision language models
Rim Assouel, Declan Campbell, Yoshua Bengio +1
To accurately process a visual scene, observers must bind features together to represent individual objects. This capacity is necessary, for instance, to distinguish an image conta…