2 papers
cs.CL2026
Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors
Yangfan Hu, Xuhan Tong, Haoyue Bai +5
Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or wh…
cs.AI2026
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…