20 citations · 20 across the 1 of their papers we have counts for
1 paper
Odalric-Ambrym Maillard, Phuong Nguyen, Ronald Ortner +1
We consider an agent interacting with an environment in a single stream of actions, observations, and rewards, with no reset. This process is not assumed to be a Markov Decision Pr…