107 citations
- Centre National de la Recherche ScientifiqueFR6 papers
- Centre de Recherche en Informatique, Signal et Automatique de LilleFR4 papers
- Google DeepMind (United Kingdom)GB4 papers
- Google (United States)US3 papers
- Centre de Recherche en InformatiqueFR2 papers
- École Centrale de LilleFR2 papers
- Meta (Israel)IL2 papers
- Sequella (United States)US2 papers
- Université de LilleFR2 papers
- Ahlia UniversityBH1 paper
- Bioengineering CenterRU1 paper
- Bioversity InternationalCO1 paper
16 papers · 1 filter
Entropy Regularized Reinforcement Learning with Cascading Networks
Riccardo Della Vecchia, Alena Shilova, Philippe Preux +1
Deep Reinforcement Learning (Deep RL) has had incredible achievements on high dimensional problems, yet its learning process remains unstable even on the simplest tasks. Deep RL us…
Soft Action Priors: Towards Robust Policy Transfer
Matheus Centa, Philippe Preux
Despite success in many challenging problems, reinforcement learning (RL) is still confronted with sample inefficiency, which can be mitigated by introducing prior knowledge to age…
When Privacy Meets Partial Information: A Refined Analysis of Differentially Private Bandits
Achraf Azize, Debabrota Basu
We study the problem of multi-armed bandits with -global Differential Privacy (DP). First, we prove the minimax and problem-dependent regret lower bounds for stochastic and line…
Optimistic PAC Reinforcement Learning: the Instance-Dependent View
Andrea Tirinzoni, Aymen Al-Marjani, Emilie Kaufmann
Optimistic algorithms have been extensively studied for regret minimization in episodic tabular MDPs, both from a minimax and an instance-dependent view. However, for the PAC RL pr…
Efficient Algorithms for Extreme Bandits
Dorian Baudry, Yoan Russac, Emilie Kaufmann
In this paper, we contribute to the Extreme Bandit problem, a variant of Multi-Armed Bandits in which the learner seeks to collect the largest possible reward. We first study the c…
Reinforcement Learning in Linear MDPs: Constant Regret and Representation Selection
Matteo Papini, Andrea Tirinzoni, Aldo Pacchiano +3
We study the role of the representation of state-action value functions in regret minimization in finite-horizon Markov Decision Processes (MDPs) with linear structure. We first de…