4 papers
The Impossibility of Eliciting Latent Knowledge
Korbinian Friedl, Francis Rhys Ward, Paul Yushin Rapoport +2
Advanced AI systems have extensive knowledge of their environments; in fact, their knowledge may (far) exceed that of their developers or users. Consequently, a desirable property…
General agents contain world models
Jonathan Richens, David Abel, Alexis Bellot +1
Are world models a necessary ingredient for flexible, goal-directed behaviour, or is model-free learning sufficient? We provide a formal answer to this question, showing that any a…
The Limits of Predicting Agents from Behaviour
Alexis Bellot, Jonathan Richens, Tom Everitt
As the complexity of AI systems and their interactions with the world increases, generating explanations for their behaviour is important for safely deploying AI. For agents, the m…
Evaluating the Goal-Directedness of Large Language Models
Tom Everitt, Cristina Garbacea, Alexis Bellot +4
To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require in…