3 papers
cs.CL2025
Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-Play
Barkavi Sundararajan, Somayajulu Sripada, Ehud Reiter
A major concern when deploying LLMs in accuracy-critical domains such as sports reporting is that the generated text may not faithfully reflect the input data. We quantify how inpu…
cs.HC2025
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
Karen Jia-Hui Li, Simone Balloccu, Ondrej Dusek +1
The increasing trust in large language models (LLMs), especially in the form of chatbots, is often undermined by the lack of their extrinsic evaluation. This holds particularly tru…
cs.HC2025
SPHERE: An Evaluation Card for Human-AI Systems
Qianou Ma, Dora Zhao, Xinran Zhao +6
In the era of Large Language Models (LLMs), establishing effective evaluation methods and standards for diverse human-AI interaction systems is increasingly challenging. To encoura…