15 papers
Information-Based Exploration via Random Features for Reinforcement Learning
Waris Radji, Odalric-Ambrym Maillard
Representation learning has enabled classical exploration strategies to be extended to deep Reinforcement Learning (RL), but often makes algorithms more complex and theoretical gua…
Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts
Udvas Das, Waris Radji, Debabrota Basu +1
We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized pref…
Pliable rejection sampling
Akram Erraqabi, Michal Valko, Alexandra Carpentier +1
Rejection sampling is a technique for sampling from difficult distributions. However, its use is limited due to a high rejection rate. Common adaptive rejection sampling methods ei…
Hierarchical Subspaces of Policies for Continual Offline Reinforcement Learning
Anthony Kobanda, Rémy Portelas, Odalric-Ambrym Maillard +1
We consider a Continual Reinforcement Learning setup, where a learning agent must continuously adapt to new tasks while retaining previously acquired skill sets, with a focus on th…
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
Anthony Kobanda, Waris Radji, Mathieu Petitbois +2
Offline Goal-Conditioned Reinforcement Learning seeks to train agents to reach specified goals from previously collected trajectories. Scaling that promises to long-horizon tasks r…
A Continual Offline Reinforcement Learning Benchmark for Navigation Tasks
Anthony Kobanda, Odalric-Ambrym Maillard, Rémy Portelas
Autonomous agents operating in domains such as robotics or video game simulations must adapt to changing tasks without forgetting about the previous ones. This process called Conti…