10 papers
The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
Kelly Cui, Nikhil Prakash, Shoval Messica +4
Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Ye…
From Activation to Specificity: Automating Counterfactual Testing of Visual Representations in the Human Brain
Yuval Golbari, Navve Wasserman, Matias Cosarinsky +5
Identifying which brain regions represent a visual concept in the human brain is a central challenge in neuroscience. Existing approaches have localized coarse functional regions (…
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
Kabir Swain, Sijie Han, Daniel Karl I. Weidele +2
Autoregressive Transformer KV caches grow linearly with context length; sliding-window caching bounds memory but discards evicted tokens entirely, so relevant evidence outside the…
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
Shaden Alshammari, Kevin Wen, Abrar Zainal +5
Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and t…
End-to-End Training for Unified Tokenization and Latent Denoising
Shivam Duggal, Xingjian Bai, Zongze Wu +5
Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer m…
Fairness Aware Reward Optimization
Ching Lam Choi, Vighnesh Subramaniam, Phillip Isola +2
Demographic skews in human preference data propagate systematic unfairness through reward models into aligned LLMs. We introduce Fairness Aware Reward Optimization (Faro), an in-pr…