3 citations · 3 across the 1 of their papers we have counts for
2 papers
cs.LG2020
Tightening Exploration in Upper Confidence Reinforcement Learning
Hippolyte Bourel, Odalric-Ambrym Maillard, Mohammad Sadegh Talebi
The upper confidence reinforcement learning (UCRL2) algorithm introduced in (Jaksch et al., 2010) is a popular method to perform regret minimization in unknown discrete Markov Deci…
cs.LG2019★ 3 cited
Model-Based Reinforcement Learning Exploiting State-Action Equivalence
Mahsa Asadi, Mohammad Sadegh Talebi, Hippolyte Bourel +1
Leveraging an equivalence property in the state-space of a Markov Decision Process (MDP) has been investigated in several studies. This paper studies equivalence structure in the r…