5 papers
Capturing Individual Human Preferences with Reward Features
André Barreto, Vincent Dumoulin, Yiran Mao +6
Reinforcement learning from human feedback usually models preferences using a reward function that does not distinguish between people. We argue that this is unlikely to be a good…
Constructing an Optimal Behavior Basis for the Option Keyboard
Lucas N. Alegre, Ana L. C. Bazzan, André Barreto +1
Multi-task reinforcement learning aims to quickly identify solutions for new tasks with minimal or no additional interaction with the environment. Generalized Policy Improvement (G…
Plasticity as the Mirror of Empowerment
David Abel, Michael Bowling, André Barreto +13
Agents are minimally entities that are influenced by their past observations and act to influence future observations. This latter capacity is captured by empowerment, which has se…
Optimizing Return Distributions with Distributional Dynamic Programming
Bernardo Ãvila Pires, Mark Rowland, Diana Borsa +6
We introduce distributional dynamic programming (DP) methods for optimizing statistical functionals of the return distribution, with standard reinforcement learning as a special ca…
Agency Is Frame-Dependent
David Abel, André Barreto, Michael Bowling +13
Agency is a system's capacity to steer outcomes toward a goal, and is a central topic of study across biology, philosophy, cognitive science, and artificial intelligence. Determini…