activity
20232026
most citedProgram-Based Strategy Induction for Reinforcement Learning

5 citations · 5 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

The Computational Basis of Confidence in Large Language Models

Dharshan Kumaran, Viorica Patraucean, Maks Ovsjanikov +2

Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated c…

cs.LG2026

ATLAS: Active Theory Learning for Automated Science

Noémi Éltető, Nathaniel D. Daw, Kimberly L. Stachenfeld +1

Advancing scientific understanding through mechanistic modeling requires posing the right experimental questions to yield maximally informative data. To automate this pursuit withi…

cs.LG2026

How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals

Dharshan Kumaran, Viorica Patraucean, Simon Osindero +2

Large language models can detect their own errors and sometimes correct them without external feedback, but the underlying mechanisms remain unknown. We investigate this through th…

cs.LG2026

Causal Evidence that Language Models use Confidence to Drive Behavior

Dharshan Kumaran, Nathaniel Daw, Simon Osindero +2

Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can…

cs.LG20245 cited

Program-Based Strategy Induction for Reinforcement Learning

Carlos G. Correa, Thomas L. Griffiths, Nathaniel D. Daw

Typical models of learning assume incremental estimation of continuously-varying decision variables like expected rewards. However, this class of models fails to capture more idios…