3 papers
cs.AI2026
Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention
Chubin Zhang, Zhenglin Wan, Xingrui Yu +5
Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a thre…
cs.LG2026
Embedding Perturbation may Better Reflect Intermediate-Step Uncertainty in LLM Reasoning
Qihao Wen, Jiahao Wang, Yang Nan +3
Large language Models (LLMs) have achieved significant breakthroughs across diverse domains; however, they can still produce unreliable or misleading outputs. For responsible LLM a…
cs.LG2026
Interpretable Probability Estimation with LLMs via Shapley Reconstruction
Yang Nan, Qihao Wen, Jiahao Wang +4
Large Language Models (LLMs) demonstrate potential to estimate the probability of uncertain events, by leveraging their extensive knowledge and reasoning capabilities. This ability…