Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
Amita Kamath, Jack Hessel, Khyathi Chandu +3
The lack of reasoning capabilities in Vision-Language Models (VLMs) has remained at the forefront of research discourse. We posit that this behavior stems from a reporting bias in…
cs.CL2026
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
Zhengyu Hu, Jieyu Zhang, Zhihan Xiong +3
Despite the remarkable success of Large Language Models (LLMs), evaluating their outputs' quality regarding preference remains a critical challenge. While existing works usually le…