4 papers · 1 filter
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
Matt MacDermott, Qiyao Wei, Rada Djoneva +1
AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as…
Measuring Goal-Directedness
Matt MacDermott, James Fox, Francesco Belardinelli +1
We define maximum entropy goal-directedness (MEG), a formal measure of goal-directedness in causal models and Markov decision processes, and give algorithms for computing it. Measu…
The Reasons that Agents Act: Intention and Instrumental Goals
Francis Rhys Ward, Matt MacDermott, Francesco Belardinelli +2
Intention is an important and challenging concept in AI. It is important because it underlies many other concepts we care about, such as agency, manipulation, legal responsibility,…
Characterising Decision Theories with Mechanised Causal Graphs
Matt MacDermott, Tom Everitt, Francesco Belardinelli
How should my own decisions affect my beliefs about the outcomes I expect to achieve? If taking a certain action makes me view myself as a certain type of person, it might affect h…