activity
20162022
most citedA unified view of entropy-regularized Markov decision processes

98 citations · 127 across the 10 of their papers we have counts for

collaborators
Showing 2020Show all

8 papers · 1 filter

cs.LG20203 cited

Hierarchical reinforcement learning for efficient exploration and transfer

Lorenzo Steccanella, Simone Totaro, Damien Allonsius +1

Sparse-reward domains are challenging for reinforcement learning algorithms since significant exploration is needed before encountering reward for the first time. Hierarchical rein…

cs.LG2020

Improved Exploration in Factored Average-Reward MDPs

Mohammad Sadegh Talebi, Anders Jonsson, Odalric-Ambrym Maillard

We consider a regret minimization task under the average-reward criterion in an unknown Factored Markov Decision Process (FMDP). More specifically, we consider an FMDP where the st…

cs.AI2020

Induction and Exploitation of Subgoal Automata for Reinforcement Learning

Daniel Furelos-Blanco, Mark Law, Anders Jonsson +2

In this paper we present ISA, an approach for learning and exploiting subgoals in episodic reinforcement learning (RL) tasks. ISA interleaves reinforcement learning with the induct…

cs.LG2020

Fast active learning for pure exploration in reinforcement learning

Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson +3

Realistic environments often provide agents with very limited feedback. When the environment is initially unknown, the feedback, in the beginning, can be completely absent, and the…

cs.LG20203 cited

Planning in Markov Decision Processes with Gap-Dependent Sample Complexity

Anders Jonsson, Emilie Kaufmann, Pierre Ménard +3

We propose MDP-GapE, a new trajectory-based Monte-Carlo Tree Search algorithm for planning in a Markov Decision Process in which transitions have a finite support. We prove an uppe…

cs.LG2020

Adaptive Reward-Free Exploration

Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues +3

Reward-free exploration is a reinforcement learning setting studied by Jin et al. (2020), who address it by running several algorithms with regret guarantees in parallel. In our wo…