44 citations · 155 across the 13 of their papers we have counts for
4 papers · 1 filter
Optimal Regret Bounds for Selecting the State Representation in Reinforcement Learning
Odalric-Ambrym Maillard, Phuong Nguyen, Ronald Ortner +1
We consider an agent interacting with an environment in a single stream of actions, observations, and rewards, with no reset. This process is not assumed to be a Markov Decision Pr…
A consistent clustering-based approach to estimating the number of change-points in highly dependent time-series
Azaden Khaleghi, Daniil Ryabko
The problem of change-point estimation is considered under a general framework where the data are generated by unknown stationary ergodic process distributions. In this context, th…
Selecting the State-Representation in Reinforcement Learning
Odalric-Ambrym Maillard, Rémi Munos, Daniil Ryabko
The problem of selecting the right state-representation in a reinforcement learning problem is considered. Several models (functions mapping past observations to a finite set) of t…
Online Regret Bounds for Undiscounted Continuous Reinforcement Learning
Ronald Ortner, Daniil Ryabko
We derive sublinear regret bounds for undiscounted reinforcement learning in continuous state space. The proposed algorithm combines state aggregation with the use of upper confide…