#confidence calibration

topicconfidence calibration

6 papers · 1 filter

cs.LG2026

Cybersecurity Detection Classification with Reasoning-enabled Language Models

Amol Khanna, Manu Nandan, Cristian Viorel Popa +10

The paper introduces a chain-of-thought reasoning classifier built on large language models to triage Windows endpoint security alerts, using a calibrated confidence estimator to i…

cs.AI2026

One Human, Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence

Cesare Zavattari, Alessandro Tommasi, Giuseppe Prencipe

The paper studies how a single human can audit a large fleet of LLM agents under a limited audit budget, analyzing how miscalibrated confidence scores and correlated errors affect…

cs.CV2026

Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers

Fexiang Liu, Shiye Wang, Qiang Qiu +1

The paper introduces Witness Evidence Portfolios (WEP), a method that analyzes the internal visual contributions of multimodal large language models to detect risky closed-form vis…

cs.CL2026

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration

Shuhao Li, Guodong Du, Anhao Zhao +3

The paper examines how supervised fine-tuning, reinforcement learning, and on‑policy distillation affect confidence estimates of large language models during chain‑of‑thought reaso…

cs.CV2026

Detector Confidence Signals Presence Rather Than Occlusion in Cluttered Manipulation

Yuanzhi He

The paper shows that confidence scores from open‑vocabulary object detectors do not reflect how much of a target object is visible, staying high even when the object is heavily occ…

cs.CL2026

Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

Jon-Paul Cacioli

The paper applies signal detection theory to separate a language model’s factual accuracy from the quality of its confidence estimates, introducing a model‑free metric (meta‑I_2r)…