3 papers
cs.AI2025
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
Kung-Hsiang Huang, Can Qin, Haoyi Qiu +4
Vision Language Models (VLMs) have achieved remarkable progress in multimodal tasks, yet they often struggle with visual arithmetic, seemingly simple capabilities like object count…
cs.CL2024
ReIFE: Re-evaluating Instruction-Following Evaluation
Yixin Liu, Kejian Shi, Alexander R. Fabbri +5
The automatic evaluation of instruction following typically involves using large language models (LLMs) to assess response quality. However, there is a lack of comprehensive evalua…
cs.CL2024
Evaluating Cultural and Social Awareness of LLM Web Agents
Haoyi Qiu, Alexander R. Fabbri, Divyansh Agarwal +4
As large language models (LLMs) expand into performing as agents for real-world applications beyond traditional NLP tasks, evaluating their robustness becomes increasingly importan…