2 papers
cs.LG2022
Approximate discounting-free policy evaluation from transient and recurrent states
Vektor Dewanto, Marcus Gallagher
In order to distinguish policies that prescribe good from bad actions in transient states, we need to evaluate the so-called bias of a policy from transient states. However, we obs…
cs.LG2020
Average-reward model-free reinforcement learning: a systematic review and literature mapping
Vektor Dewanto, George Dunn, Ali Eshragh +2
Reinforcement learning is important part of artificial intelligence. In this paper, we review model-free reinforcement learning that utilizes the average reward optimality criterio…