activity
20242026
most citedMAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space

1 citations · 1 across the 2 of their papers we have counts for

collaborators

8 papers

cs.CV20261 cited

MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space

Santiago Galella, Pamela Osuna-Vargas, Maren Wehrheim +3

Modern vision models achieve strong performance on standard benchmarks, yet their aggregate accuracy reveals little about which scene properties drive their predictions. Existing r…

cs.CV2026

Mechanisms of Object Localization in Vision-Language Models

Timothy Schaumlöffel, Martina G. Vilas, Gemma Roig

Visually-grounded language models (VLMs) are highly effective in linking visual and textual information, yet they often struggle with basic classification and localization tasks. W…

cs.CV2026

Contextual inference from single objects in Vision-Language models

Martina G. Vilas, Timothy Schaumlöffel, Gemma Roig

How much scene context a single object carries is a well-studied question in human scene perception, yet how this capacity is organized in vision-language models (VLMs) remains poo…

cs.AI2025

Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning

Martina G. Vilas, Safoora Yousefi, Besmira Nushi +2

Reasoning models improve their problem-solving ability through inference-time scaling, allocating more compute via longer token budgets. Identifying which reasoning traces are like…

cs.CV2025

Net2Brain: A Toolbox to compare artificial vision models with human brain responses

Domenic Bersch, Kshitij Dwivedi, Martina Vilas +2

We introduce Net2Brain, a graphical and command-line user interface toolbox for comparing the representational spaces of artificial deep neural networks (DNNs) and human brain reco…

cs.CV2025

FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks

Mahadev Prasad Panda, Matteo Tiezzi, Martina Vilas +3

Explainability in artificial intelligence (XAI) remains a crucial aspect for fostering trust and understanding in machine learning models. Current visual explanation techniques, su…