Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Bootstrapping Expectiles in Reinforcement Learning
Pierre Clavier, Emmanuel Rachelson, Erwan Le Pennec +1
Many classic Reinforcement Learning (RL) algorithms rely on a Bellman operator, which involves an expectation over the next states, leading to the concept of bootstrapping. To intr…
cs.LG2024
Towards Minimax Optimality of Model-based Robust Reinforcement Learning
Pierre Clavier, Erwan Le Pennec, Matthieu Geist
We study the sample complexity of obtaining an -optimal policy in \emph{Robust} discounted Markov Decision Processes (RMDPs), given only access to a generative model of the nom…