1 paper · 1 filter
Hamid Reza Maei
We present the first class of policy-gradient algorithms that work with both state-value and policy function-approximation, and are guaranteed to converge under off-policy training…