1 paper
Ku Onoda, Paavo Parmas, Manato Yaguchi +1
In policy gradient reinforcement learning, access to a differentiable model enables 1st-order gradient estimation that accelerates learning compared to relying solely on derivative…