activity
20242026
most citedAdaptive Length Image Tokenization via Recurrent Allocation

1 citations · 1 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

From Activation to Specificity: Automating Counterfactual Testing of Visual Representations in the Human Brain

Yuval Golbari, Navve Wasserman, Matias Cosarinsky +5

Identifying which brain regions represent a visual concept in the human brain is a central challenge in neuroscience. Existing approaches have localized coarse functional regions (…

cs.CV2026

End-to-End Training for Unified Tokenization and Latent Denoising

Shivam Duggal, Xingjian Bai, Zongze Wu +5

Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer m…

cs.CV2026

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

Kelly Cui, Nikhil Prakash, Shoval Messica +4

Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Ye…

cs.CV2025

Separating Knowledge and Perception with Procedural Data

Adrián Rodríguez-Muñoz, Manel Baradad, Phillip Isola +1

We train representation models with procedural data only, and apply them on visual similarity, classification, and semantic segmentation tasks without further training by using vis…

cs.CV2025

Single-pass Adaptive Image Tokenization for Minimum Program Search

Shivam Duggal, Sanghyun Byun, William T. Freeman +2

According to Algorithmic Information Theory (AIT) -- Intelligent representations compress data into the shortest possible program that can reconstruct its content, exhibiting low K…

cs.CV2025

Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation

Shivam Duggal, Yushi Hu, Oscar Michel +7

Despite the unprecedented progress in the field of 3D generation, current systems still often fail to produce high-quality 3D assets that are visually appealing and geometrically a…