works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.LG2026

The Computational Basis of Confidence in Large Language Models

Dharshan Kumaran, Viorica Patraucean, Maks Ovsjanikov +3

The paper investigates what the confidence signal in large language models actually represents, showing that answer logits often act as monotonic readouts of a latent decision vari…

cs.LG2026

Reported Confidence in LLMs Tracks Commitment More Than Correctness

Dharshan Kumaran

Confidence is an estimate of the probability that a chosen answer is correct. Verbal confidence reports are widely used as uncertainty measures in large language models, but whethe…

cs.LG2026

Causal Evidence that Language Models use Confidence to Drive Behavior

Dharshan Kumaran, Nathaniel Daw, Simon Osindero +2

Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can…

cs.CL2026

How do LLMs Compute Verbal Confidence

Dharshan Kumaran, Arthur Conmy, Federico Barbero +3

Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs in…

cs.LG2026

How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals

Dharshan Kumaran, Viorica Patraucean, Simon Osindero +2

Large language models can detect their own errors and sometimes correct them without external feedback, but the underlying mechanisms remain unknown. We investigate this through th…

cs.LG2025

How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models

Dharshan Kumaran, Stephen M Fleming, Larisa Markeeva +8

Large language models (LLMs) exhibit strikingly conflicting behaviors: they can appear steadfastly overconfident in their initial answers whilst at the same time being prone to exc…