4 papers
An Information-Theoretic Definition for Open-Ended Learning
Wanqiao Xu, Yifan Zhu, Benjamin Van Roy
A growing body of work points to the great promise of AI systems that can continually expand their capabilities as they operate in an open-ended environment. But yet there is no co…
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Jonathan Colaço Carr, Jonathan Colaço Carr, Prakash Panangaden +2
Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to…
Efficient Exploration at Scale
Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla +5
We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates rew…
Capacity-Constrained Continual Learning
Zheng Wen, Doina Precup, Benjamin Van Roy +1
Any agents we can possibly build are subject to capacity constraints, as memory and compute resources are inherently finite. However, comparatively little attention has been dedica…