activity
20242026
collaborators

15 papers

cs.LG2026

Information-Based Exploration via Random Features for Reinforcement Learning

Waris Radji, Odalric-Ambrym Maillard

Representation learning has enabled classical exploration strategies to be extended to deep Reinforcement Learning (RL), but often makes algorithms more complex and theoretical gua…

cs.LG2026

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

Udvas Das, Waris Radji, Debabrota Basu +1

We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized pref…

stat.ML2026

Pliable rejection sampling

Akram Erraqabi, Michal Valko, Alexandra Carpentier +1

Rejection sampling is a technique for sampling from difficult distributions. However, its use is limited due to a high rejection rate. Common adaptive rejection sampling methods ei…

cs.LG2026

Hierarchical Subspaces of Policies for Continual Offline Reinforcement Learning

Anthony Kobanda, Rémy Portelas, Odalric-Ambrym Maillard +1

We consider a Continual Reinforcement Learning setup, where a learning agent must continuously adapt to new tasks while retaining previously acquired skill sets, with a focus on th…

cs.LG2026

Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning

Anthony Kobanda, Waris Radji, Mathieu Petitbois +2

Offline Goal-Conditioned Reinforcement Learning seeks to train agents to reach specified goals from previously collected trajectories. Scaling that promises to long-horizon tasks r…

cs.LG2026

A Continual Offline Reinforcement Learning Benchmark for Navigation Tasks

Anthony Kobanda, Odalric-Ambrym Maillard, Rémy Portelas

Autonomous agents operating in domains such as robotics or video game simulations must adapt to changing tasks without forgetting about the previous ones. This process called Conti…