90 citations · 194 across the 7 of their papers we have counts for
Showing 2013Show all
2 papers · 1 filter
math.OC2013★ 2 cited
Tight Performance Bounds for Approximate Modified Policy Iteration with Non-Stationary Policies
Boris Lesner, Bruno Scherrer
We consider approximate dynamic programming for the infinite-horizon stationary -discounted optimal control problem formalized by Markov Decision Processes. While in the exact c…
cs.AI2013★ 39 cited
Off-policy Learning with Eligibility Traces: A Survey
Matthieu Geist, Bruno Scherrer
In the framework of Markov Decision Processes, off-policy learning, that is the problem of learning a linear approximation of the value function of some fixed policy from one traje…