11 papers
The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
Kelly Cui, Nikhil Prakash, Shoval Messica +4
Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Ye…
BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain
Navve Wasserman, Matias Cosarinsky, Yuval Golbari +4
Understanding how the human brain represents visual concepts, and in which brain regions these representations are encoded, remains a long-standing challenge. Decades of work have…
Text-to-Image Models Need Less from Text Encoders Than You Think
Nurit Spingarn, Noa Cohen, Tamar Rott Shaham +1
Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that condition the image generation proc…
Vision-Language Binding in In-Context Image Generation
Chris Ge, Rohit Gandikota, Antonio Torralba +1
In-context image generation models such as FLUX.2 take a text prompt and an optional reference image as visual conditioning for the output. Internally, all three inputs -- text, re…
From Activation to Specificity: Automating Counterfactual Testing of Visual Representations in the Human Brain
Yuval Golbari, Navve Wasserman, Matias Cosarinsky +5
Identifying which brain regions represent a visual concept in the human brain is a central challenge in neuroscience. Existing approaches have localized coarse functional regions (…
Letting the neural code speak: Automated characterization of monkey visual neurons through human language
Vedang Lad, Katrin Franke, Tamar Rott Shaham +4
Understanding what individual neurons encode is a core question in neuroscience. In primary visual cortex (V1), mathematical models (e.g., Gabor functions) capture neural selectivi…