5 papers · 1 filter
Prior Diffusiveness and Regret in the Linear-Gaussian Bandit
Yifan Zhu, John C. Duchi, Benjamin Van Roy
We prove that Thompson sampling exhibits Bayesian regret in the linear-Gaussian bandit with a pri…
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
Henrik Marklund, Benjamin Van Roy
As AI agents generate increasingly sophisticated behaviors, manually encoding human preferences to guide these agents becomes more challenging. To address this, it has been suggest…
Aligning AI Agents via Information-Directed Sampling
Hong Jun Jeon, Benjamin Van Roy
The staggering feats of AI systems have brought to attention the topic of AI Alignment: aligning a "superintelligent" AI agent's actions with humanity's interests. Many existing fr…
The Need for a Big World Simulator: A Scientific Challenge for Continual Learning
Saurabh Kumar, Hong Jun Jeon, Alex Lewandowski +1
The "small agent, big world" frame offers a conceptual view that motivates the need for continual learning. The idea is that a small agent operating in a much bigger world cannot s…
Information-Theoretic Foundations for Neural Scaling Laws
Hong Jun Jeon, Benjamin Van Roy
Neural scaling laws aim to characterize how out-of-sample error behaves as a function of model and training dataset size. Such scaling laws guide allocation of a computational reso…