2 papers
cs.CL2026
PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
Khizar Hussain, Murat Kantarcioglu
Large language models (LLMs) hallucinate with confidence: their outputs can be fluent, authoritative, and simply wrong. In medical, legal, and scientific applications this failure…
cs.CL2026
Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
Khizar Hussain, Bradley A. Malin, Zhijun Yin +2
As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safety. However, state-of-the-art L…