Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty
Ali Şenol, H. Russell Bernard, Huan Liu
Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: do they know what they don't know? To answer this question, we borrow t…
cs.AI2026
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
Ali Şenol, Garima Agrawal, Huan Liu
Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason,…