6 papers
Hues and Cues: Human vs. CLIP
Nuria Alabau-Bosque, Jorge Vila-Tomás, Paula Daudén-Oliver +4
Playing games is inherently human, and a lot of games are created to challenge different human characteristics. However, these tasks are often left out when evaluating the human-li…
Do Vision Transformers See Like Humans? Evaluating their Perceptual Alignment
Pablo Hernández-Cámara, Jose Manuel Jaén-Lorites, Jorge Vila-Tomás +2
Vision Transformers (ViTs) achieve remarkable performance in image recognition tasks, yet their alignment with human perception remains largely unexplored. This study systematicall…
Contrast Sensitivity in Multimodal Large Language Models: A Psychophysics-Inspired Evaluation
Pablo Hernández-Cámara, Alexandra Gomez-Villa, Jose Manuel Jaén-Lorites +3
Understanding how Multimodal Large Language Models (MLLMs) process low-level visual features is critical for evaluating their perceptual abilities and has not been systematically c…
On the dynamic evolution of CLIP texture-shape bias and its relationship to human alignment and model robustness
Pablo Hernández-Cámara, Jose Manuel Jaén-Lorites, Alexandra Gómez-Villa +3
Contrastive language-image models such as CLIP have demonstrated remarkable generalization capabilities. However, how their internal visual representations evolve during training a…
A Turing Test for Artificial Nets devoted to model Human Vision
Jorge Vila-Tomás, Pablo Hernández-Cámara, Qiang Li +2
In our invited talk at the AI Evaluation Workshop of the University of Bristol back in June 2022 we argued that, despite claims about successful modeling of the visual brain using…
Parametric PerceptNet: A bio-inspired deep-net trained for Image Quality Assessment
Jorge Vila-Tomás, Pablo Hernández-Cámara, Valero Laparra +1
Human vision models are at the core of image processing. For instance, classical approaches to the problem of image quality are based on models that include knowledge about human v…