3 papers
cs.CL2026
MASH: Modeling Abstention via Selective Help-Seeking
Mustafa Omer Gul, Claire Cardie, Tanya Goyal
LLMs cannot reliably recognize their parametric knowledge boundaries and often hallucinate answers to outside-of-boundary questions. In this paper, we introduce MASH (Modeling Abst…
cs.HC2024
Challenges in Trustworthy Human Evaluation of Chatbots
Wenting Zhao, Alexander M. Rush, Tanya Goyal
Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available b…
cs.CL2024
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
Wenting Zhao, Tanya Goyal, Yu Ying Chiu +8
While hallucinations of large language models (LLMs) prevail as a major challenge, existing evaluation benchmarks on factuality do not cover the diverse domains of knowledge that t…