activity
20242026
collaborators

17 papers

cs.CY2026

AI Value Alignment for Evolving Social Norms

Nenad Tomašev, Matija Franklin, Simon Osindero

AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop…

cs.LG2026

Causal Evidence that Language Models use Confidence to Drive Behavior

Dharshan Kumaran, Nathaniel Daw, Simon Osindero +2

Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can…

cs.CL2026

How do LLMs Compute Verbal Confidence

Dharshan Kumaran, Arthur Conmy, Federico Barbero +3

Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs in…

cs.AI2026

Distributional AGI Safety

Nenad Tomašev, Matija Franklin, Julian Jacobs +2

AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithi…

cs.AI2026

MaD Physics: Evaluating information seeking under constraints in physical environments

Moksh Jain, Mehdi Bennani, Johannes Bausch +4

Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of measurements due to physical an…

cs.LG2026

How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals

Dharshan Kumaran, Viorica Patraucean, Simon Osindero +2

Large language models can detect their own errors and sometimes correct them without external feedback, but the underlying mechanisms remain unknown. We investigate this through th…