paper

Evaluating Vision-Language Models for Image Quality Assessment using Psychophysical Data

arXiv:2603.24578

Abstract

Psychophysical experiments remain the most reliable approach for perceptual image quality assessment (IQA), yet their cost and limited scalability motivate automated alternatives. This paper investigates whether Vision-Language Models (VLMs) can assist in assessing perceived image appearance and quality. We introduce a psychophysics-inspired framework to probe VLM perceptual sensitivity through controlled pairwise image comparisons of contrast, colorfulness, and overall preference. Six VLMs (four proprietary and two open-weight models) are compared against psychophysical data. Results reveal strong attribute-dependent variability: Claude exhibits the highest internal consistency, whereas GPT achieves the strongest agreement for overall preference. Claude and Qwen show the strongest alignment for colorfulness, while Qwen performs best for contrast. However, no model consistently matches human perception across all attributes. High self-consistency does not necessarily imply perceptual validity, and VLM--human agreement generally improves when perceptual differences among renderings are more pronounced.

Accepted at CIC'34