1 paper
Mengmeng Li, Daniel Kuhn, Tobias Sutter
We study offline reinforcement learning problems with a long-run average reward objective. The state-action pairs generated by any fixed behavioral policy thus follow a Markov chai…