activity
20102026
most citedModel-Based Reinforcement Learning Exploiting State-Action Equivalence

3 citations · 3 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2024

Provably Efficient Exploration in Reward Machines with Low Regret

Hippolyte Bourel, Anders Jonsson, Odalric-Ambrym Maillard +2

We study reinforcement learning (RL) for decision processes with non-Markovian reward, in which high-level knowledge of the task in the form of reward machines is available to the…

cs.LG2024

Near-Optimal Reinforcement Learning with Shuffle Differential Privacy

Shaojie Bai, Mohammad Sadegh Talebi, Chengcheng Zhao +2

Reinforcement learning (RL) is a powerful tool for sequential decision-making, but its application is often hindered by privacy concerns arising from its interaction data. This cha…

cs.LG2024

Tractable Offline Learning of Regular Decision Processes

Ahana Deb, Roberto Cipollone, Anders Jonsson +2

This work studies offline Reinforcement Learning (RL) in a class of non-Markovian environments called Regular Decision Processes (RDPs). In RDPs, the unknown dependency of future o…

cs.LG2020

Improved Exploration in Factored Average-Reward MDPs

Mohammad Sadegh Talebi, Anders Jonsson, Odalric-Ambrym Maillard

We consider a regret minimization task under the average-reward criterion in an unknown Factored Markov Decision Process (FMDP). More specifically, we consider an FMDP where the st…

cs.LG2020

Tightening Exploration in Upper Confidence Reinforcement Learning

Hippolyte Bourel, Odalric-Ambrym Maillard, Mohammad Sadegh Talebi

The upper confidence reinforcement learning (UCRL2) algorithm introduced in (Jaksch et al., 2010) is a popular method to perform regret minimization in unknown discrete Markov Deci…

cs.LG20193 cited

Model-Based Reinforcement Learning Exploiting State-Action Equivalence

Mahsa Asadi, Mohammad Sadegh Talebi, Hippolyte Bourel +1

Leveraging an equivalence property in the state-space of a Markov Decision Process (MDP) has been investigated in several studies. This paper studies equivalence structure in the r…