14 papers
Human-AI Perceptual Alignment by Playing Hues and Cues
Nuria Alabau-Bosque, Jorge Vila-Tomás, Paula Daudén-Oliver +3
Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks that overlook fine-grained sema…
Parameter-Efficient Architectural Modifications for Translation-Invariant CNNs
Nuria Alabau-Bosque, Jorge Vila-Tomas, Paula Dauden-Oliver +2
Convolutional Neural Networks (CNNs) are widely assumed to be translation-invariant, yet standard architectures exhibit a startling fragility: even a single-pixel shift can drastic…
Image Segmentation via Divisive Normalization: dealing with environmental diversity
Pablo Hernández-Cámara, Jorge Vila-Tomás, Paula Dauden-Oliver +3
Autonomous driving is a challenging scenario for image segmentation due to the presence of uncontrolled environmental conditions and the eventually catastrophic consequences of fai…
On the RAID dataset of perceptual responses: analysis and statistical causes
Paula Daudén-Oliver, David Agost-Beltran, Emilio Sansano-Sansano +4
This work analyzes the RAID dataset to evaluate human responses to affine image distortions, including rotation, translation, scaling, and Gaussian noise. Using Mean Squared Error…
On the dynamic evolution of CLIP texture-shape bias and its relationship to human alignment and model robustness
Pablo Hernández-Cámara, Jose Manuel Jaén-Lorites, Alexandra Gómez-Villa +3
Contrastive language-image models such as CLIP have demonstrated remarkable generalization capabilities. However, how their internal visual representations evolve during training a…
Contrast Sensitivity in Multimodal Large Language Models: A Psychophysics-Inspired Evaluation
Pablo Hernández-Cámara, Alexandra Gomez-Villa, Jose Manuel Jaén-Lorites +3
Understanding how Multimodal Large Language Models (MLLMs) process low-level visual features is critical for evaluating their perceptual abilities and has not been systematically c…