6 papers
MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
Santiago Galella, Pamela Osuna-Vargas, Maren Wehrheim +3
Modern vision models achieve strong performance on standard benchmarks, yet their aggregate accuracy reveals little about which scene properties drive their predictions. Existing r…
Mechanisms of Object Localization in Vision-Language Models
Timothy Schaumlöffel, Martina G. Vilas, Gemma Roig
Visually-grounded language models (VLMs) are highly effective in linking visual and textual information, yet they often struggle with basic classification and localization tasks. W…
Evaluation of Randomization through Style Transfer for Enhanced Domain Generalization
Dustin Eisenhardt, Timothy Schaumlöffel, Alperen Kantarci +1
Deep learning models for computer vision often suffer from poor generalization when deployed in real-world settings, especially when trained on synthetic data due to the well-known…
Temporal Slowness in Central Vision Drives Semantic Object Learning
Timothy Schaumlöffel, Arthur Aubret, Gemma Roig +1
Humans acquire semantic object representations from egocentric visual streams with minimal supervision, but the underlying mechanisms remain unclear. Importantly, the visual system…
Contextual inference from single objects in Vision-Language models
Martina G. Vilas, Timothy Schaumlöffel, Gemma Roig
How much scene context a single object carries is a well-studied question in human scene perception, yet how this capacity is organized in vision-language models (VLMs) remains poo…
Human Gaze Boosts Object-Centered Representation Learning
Timothy Schaumlöffel, Arthur Aubret, Gemma Roig +1
Recent self-supervised learning (SSL) models trained on human-like egocentric visual inputs substantially underperform on image recognition tasks compared to humans. These models t…