#confidence calibration
6 papers · 1 filter
Cybersecurity Detection Classification with Reasoning-enabled Language Models
Amol Khanna, Manu Nandan, Cristian Viorel Popa +10
The paper introduces a chain-of-thought reasoning classifier built on large language models to triage Windows endpoint security alerts, using a calibrated confidence estimator to i…
One Human, Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence
Cesare Zavattari, Alessandro Tommasi, Giuseppe Prencipe
The paper studies how a single human can audit a large fleet of LLM agents under a limited audit budget, analyzing how miscalibrated confidence scores and correlated errors affect…
Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers
Fexiang Liu, Shiye Wang, Qiang Qiu +1
The paper introduces Witness Evidence Portfolios (WEP), a method that analyzes the internal visual contributions of multimodal large language models to detect risky closed-form vis…
Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration
Shuhao Li, Guodong Du, Anhao Zhao +3
The paper examines how supervised fine-tuning, reinforcement learning, and on‑policy distillation affect confidence estimates of large language models during chain‑of‑thought reaso…
Detector Confidence Signals Presence Rather Than Occlusion in Cluttered Manipulation
Yuanzhi He
The paper shows that confidence scores from open‑vocabulary object detectors do not reflect how much of a target object is visible, staying high even when the object is heavily occ…
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
Jon-Paul Cacioli
The paper applies signal detection theory to separate a language model’s factual accuracy from the quality of its confidence estimates, introducing a model‑free metric (meta‑I_2r)…