collaborators

7 papers

cs.LG2026

An Information-Theoretic Definition for Open-Ended Learning

Wanqiao Xu, Yifan Zhu, Benjamin Van Roy

A growing body of work points to the great promise of AI systems that can continually expand their capabilities as they operate in an open-ended environment. But yet there is no co…

cs.LG2026

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems

Jonathan Colaço Carr, Prakash Panangaden, Doina Precup +1

Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to…

cs.LG2026

Efficient Exploration at Scale

Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla +5

We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates rew…

cs.LG2026

Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

Yifan Zhu, John C. Duchi, Benjamin Van Roy

We prove that Thompson sampling exhibits Bayesian regret in the linear-Gaussian bandit with a prior d…

cs.LG2024

The Need for a Big World Simulator: A Scientific Challenge for Continual Learning

Saurabh Kumar, Hong Jun Jeon, Alex Lewandowski +1

The "small agent, big world" frame offers a conceptual view that motivates the need for continual learning. The idea is that a small agent operating in a much bigger world cannot s…

stat.ML2024

Information-Theoretic Foundations for Machine Learning

Hong Jun Jeon, Benjamin Van Roy

The progress of machine learning over the past decade is undeniable. In retrospect, it is both remarkable and unsettling that this progress was achievable with little to no rigorou…