8 papers
Information-Based Exploration via Random Features for Reinforcement Learning
Waris Radji, Odalric-Ambrym Maillard
Representation learning has enabled classical exploration strategies to be extended to deep Reinforcement Learning (RL), but often makes algorithms more complex and theoretical gua…
Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX
Waris Radji, Thomas Michel, Hector Piteau
Reinforcement learning (RL) research requires diverse, challenging environments that are both tractable and scalable. While modern video games may offer rich dynamics, they are com…
Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts
Udvas Das, Waris Radji, Debabrota Basu +1
We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized pref…
Intrinsic-Energy Joint Embedding Predictive Architectures Induce Quasimetric Spaces
Anthony Kobanda, Waris Radji
Joint-Embedding Predictive Architectures (JEPAs) aim to learn representations by predicting target embeddings from context embeddings, inducing a scalar compatibility energy in a l…
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
Anthony Kobanda, Waris Radji, Mathieu Petitbois +2
Offline Goal-Conditioned Reinforcement Learning seeks to train agents to reach specified goals from previously collected trajectories. Scaling that promises to long-horizon tasks r…
How Hard is it to Confuse a World Model?
Waris Radji, Odalric-Ambrym Maillard
In reinforcement learning (RL) theory, the concept of most confusing instances is central to establishing regret lower bounds, that is, the minimal exploration needed to solve a pr…