2 papers
cs.CL2026
Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
Nathan Mao, Varun Kaushik, Shreya Shivkumar +3
Large Language Models (LLMs) often hallucinate, generating nonsensical or false information that can be especially harmful in sensitive fields such as medicine or law. To study thi…
cs.AI2025
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
Abhishek More, Anthony Zhang, Nicole Bonilla +4
Chain-of-thought (CoT) prompting enables Large Language Models to solve complex problems, but deploying these models safely requires reliable confidence estimates, a capability whe…