6 citations · 6 across the 1 of their papers we have counts for
4 papers · 1 filter
Non-Stationary Bandit Learning via Predictive Sampling
Yueyang Liu, Xu Kuang, Benjamin Van Roy
Thompson sampling has proven effective across a wide range of stationary bandit environments. However, as we demonstrate in this paper, it can perform poorly when applied to non-st…
Posterior Sampling for Continuing Environments
Wanqiao Xu, Shi Dong, Benjamin Van Roy
We develop an extension of posterior sampling for reinforcement learning (PSRL) that is suited for a continuing agent-environment interface and integrates naturally into agent desi…
Continual Learning as Computationally Constrained Reinforcement Learning
Saurabh Kumar, Henrik Marklund, Ashish Rao +4
An agent that efficiently accumulates knowledge to develop increasingly sophisticated skills over a long lifetime could advance the frontier of artificial intelligence capabilities…
Maintaining Plasticity in Continual Learning via Regenerative Regularization
Saurabh Kumar, Henrik Marklund, Benjamin Van Roy
In continual learning, plasticity refers to the ability of an agent to quickly adapt to new information. Neural networks are known to lose plasticity when processing non-stationary…