10 papers · 1 filter
An Information-Theoretic Definition for Open-Ended Learning
Wanqiao Xu, Yifan Zhu, Benjamin Van Roy
A growing body of work points to the great promise of AI systems that can continually expand their capabilities as they operate in an open-ended environment. But yet there is no co…
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Jonathan Colaço Carr, Prakash Panangaden, Doina Precup +1
Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to…
Efficient Exploration at Scale
Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla +5
We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates rew…
Prior Diffusiveness and Regret in the Linear-Gaussian Bandit
Yifan Zhu, John C. Duchi, Benjamin Van Roy
We prove that Thompson sampling exhibits Bayesian regret in the linear-Gaussian bandit with a prior d…
The Need for a Big World Simulator: A Scientific Challenge for Continual Learning
Saurabh Kumar, Hong Jun Jeon, Alex Lewandowski +1
The "small agent, big world" frame offers a conceptual view that motivates the need for continual learning. The idea is that a small agent operating in a much bigger world cannot s…
Information-Theoretic Foundations for Neural Scaling Laws
Hong Jun Jeon, Benjamin Van Roy
Neural scaling laws aim to characterize how out-of-sample error behaves as a function of model and training dataset size. Such scaling laws guide allocation of a computational reso…