From the 1 of 13 linked papers with an AI index.
13 papers
The Computational Basis of Confidence in Large Language Models
Dharshan Kumaran, Viorica Patraucean, Maks Ovsjanikov +3
The paper investigates what the confidence signal in large language models actually represents, showing that answer logits often act as monotonic readouts of a latent decision vari…
Causal Evidence that Language Models use Confidence to Drive Behavior
Dharshan Kumaran, Nathaniel Daw, Simon Osindero +2
Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can…
How do LLMs Compute Verbal Confidence
Dharshan Kumaran, Arthur Conmy, Federico Barbero +3
Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs in…
How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
Dharshan Kumaran, Viorica Patraucean, Simon Osindero +2
Large language models can detect their own errors and sometimes correct them without external feedback, but the underlying mechanisms remain unknown. We investigate this through th…
Leveraging Classical Algorithms for Graph Neural Networks
Jason Wu, Petar VeliÄkoviÄ
Neural networks excel at processing unstructured data but often fail to generalise out-of-distribution, whereas classical algorithms guarantee correctness but lack flexibility. We…
Extracting alignment data in open models
Federico Barbero, Xiangming Gu, Christopher A. Choquette-Choo +6
In this work, we show that it is possible to extract significant amounts of alignment training data from a post-trained model -- useful to steer the model to improve certain capabi…