Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When Vision-Language Models Judge Without Seeing: Exposing Informativeness Bias
Xiaohan Zou, Roshan Sridhar, Mohammadtaher Safarzadeh +1
The reliability of VLM-as-a-Judge is critical for the automatic evaluation of vision-language models (VLMs). Despite recent progress, our analysis reveals that VLM-as-a-Judge often…
cs.AI2026
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
Yongkang Du, Xiaohan Zou, Minhao Cheng +1
Analogical reasoning tests a fundamental aspect of human cognition: mapping the relation from one pair of objects to another. Existing evaluations of this ability in multimodal lar…