7 papers
Attention Alignment Between Humans and Vision-Language Models
Isaac R. Christian, Udith Haputhanthrige, Hanna Hornfeld +4
Visual perception depends on top-down goals and bottom-up sensory mechanisms. Vision-language models implement both, allowing us to treat each component as a separable hypothesis a…
Binding Visual Features Point by Point
Udith Haputhanthri, Declan Campbell, Rim Assouel +2
Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, including many tasks that are relat…
Understanding Task Representations in Neural Networks via Bayesian Ablation
Andrew Nam, Declan Campbell, Thomas Griffiths +2
Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging…
Context Structure Reshapes the Representational Geometry of Language Models
Eghbal A. Hosseini, Yuxuan Li, Yasaman Bahri +2
Large Language Models (LLMs) have been shown to organize the representations of input sequences into straighter neural trajectories in their deep layers, which has been hypothesize…
Visual symbolic mechanisms: Emergent symbol processing in vision language models
Rim Assouel, Declan Campbell, Yoshua Bengio +1
To accurately process a visual scene, observers must bind features together to represent individual objects. This capacity is necessary, for instance, to distinguish an image conta…
Just-in-time and distributed task representations in language models
Yuxuan Li, Declan Campbell, Stephanie C. Y. Chan +1
Many of language models' impressive capabilities originate from their in-context learning: based on instructions or examples, they can infer and perform new tasks without weight up…