1 paper
Yuhua Zhu, Yutong Ren
In this paper, we develop an off-policy method for continuous-time reinforcement learning (CTRL), where the system dynamics are governed by an unknown stochastic differential equat…