3 papers
cs.LG2026
Non-Stationary Bandit Learning via Predictive Sampling
Yueyang Liu, Xu Kuang, Benjamin Van Roy
Thompson sampling has proven effective across a wide range of stationary bandit environments. However, as we demonstrate in this paper, it can perform poorly when applied to non-st…
cs.LG2025
Posterior Sampling for Continuing Environments
Wanqiao Xu, Shi Dong, Benjamin Van Roy
We develop an extension of posterior sampling for reinforcement learning (PSRL) that is suited for a continuing agent-environment interface and integrates naturally into agent desi…
cs.LG2025
Continual Learning as Computationally Constrained Reinforcement Learning
Saurabh Kumar, Henrik Marklund, Ashish Rao +4
An agent that efficiently accumulates knowledge to develop increasingly sophisticated skills over a long lifetime could advance the frontier of artificial intelligence capabilities…