1 citations · 1 across the 11 of their papers we have counts for
7 papers · 1 filter
From Activation to Specificity: Automating Counterfactual Testing of Visual Representations in the Human Brain
Yuval Golbari, Navve Wasserman, Matias Cosarinsky +5
Identifying which brain regions represent a visual concept in the human brain is a central challenge in neuroscience. Existing approaches have localized coarse functional regions (…
End-to-End Training for Unified Tokenization and Latent Denoising
Shivam Duggal, Xingjian Bai, Zongze Wu +5
Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer m…
The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
Kelly Cui, Nikhil Prakash, Shoval Messica +4
Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Ye…
Separating Knowledge and Perception with Procedural Data
Adrián Rodríguez-Muñoz, Manel Baradad, Phillip Isola +1
We train representation models with procedural data only, and apply them on visual similarity, classification, and semantic segmentation tasks without further training by using vis…
Single-pass Adaptive Image Tokenization for Minimum Program Search
Shivam Duggal, Sanghyun Byun, William T. Freeman +2
According to Algorithmic Information Theory (AIT) -- Intelligent representations compress data into the shortest possible program that can reconstruct its content, exhibiting low K…
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
Shivam Duggal, Yushi Hu, Oscar Michel +7
Despite the unprecedented progress in the field of 3D generation, current systems still often fail to produce high-quality 3D assets that are visually appealing and geometrically a…