7 papers
An Information-Theoretic Definition for Open-Ended Learning
Wanqiao Xu, Yifan Zhu, Benjamin Van Roy
A growing body of work points to the great promise of AI systems that can continually expand their capabilities as they operate in an open-ended environment. But yet there is no co…
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Jonathan Colaço Carr, Prakash Panangaden, Doina Precup +1
Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to…
Efficient Exploration at Scale
Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla +5
We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates rew…
Prior Diffusiveness and Regret in the Linear-Gaussian Bandit
Yifan Zhu, John C. Duchi, Benjamin Van Roy
We prove that Thompson sampling exhibits Bayesian regret in the linear-Gaussian bandit with a prior d…
The Need for a Big World Simulator: A Scientific Challenge for Continual Learning
Saurabh Kumar, Hong Jun Jeon, Alex Lewandowski +1
The "small agent, big world" frame offers a conceptual view that motivates the need for continual learning. The idea is that a small agent operating in a much bigger world cannot s…
Information-Theoretic Foundations for Machine Learning
Hong Jun Jeon, Benjamin Van Roy
The progress of machine learning over the past decade is undeniable. In retrospect, it is both remarkable and unsettling that this progress was achievable with little to no rigorou…