Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs
arXiv:2610.05463 · doi:10.1007/s42113-026-00333-4
Abstract
Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both groups showed efficient performance in feature searches, most clearly when the target had a unique color, but performance degradation in conjunction searches as set sizes increased. Additionally, we found strong correlations between human and MLLM error rates (), which suggests that MLLMs are sensitive to similar objective complexities, such as stimulus heterogeneity. However, differences were found as well: whereas humans invested extra search time to respond accurately on target-absent trials, MLLMs exhibited extreme present/absent response biases in complex searches. We conclude that MLLMs replicate high-level human performance signatures, yet their underlying computations differ significantly.
References in corpus (13)
- A Survey on Multimodal Large Language Models
- Gemma 3 Technical Report
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- An Introduction to Vision-Language Modeling
- Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
- Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
- Attention Alignment Between Humans and Vision-Language Models
- Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms
- Exploring Perceptual Limitation of Multimodal Large Language Models
- From Cognition to Computation: A Comparative Review of Human Attention and Transformer Architectures
- I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
- Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
- The Geometry of Representational Failures in Vision Language Models