3 papers
cs.SE2026
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
Shashwat Pandey, Satwik Pandey, Suresh Raghu
Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device models now ship to hundreds of millions of devices with no server-side moderatio…
cs.AI2026
Proper Scoring Rules for Agentic Uncertainty Quantification
Suresh Raghu, Satwik Pandey, Shashwat Pandey
Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulness with probabilistic truthf…
cs.AI2026
SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio
Satwik Pandey, Suresh Raghu, Shashwat Pandey
Uncertainty estimation for reasoning language models remains difficult to deploy in practice: sampling-based methods are computationally expensive, while common single-pass proxies…