6 papers
A Robust Model-Based Approach for Continuous-Time Policy Evaluation with Unknown Lévy Process Dynamics
Qihao Ye, Xiaochuan Tian, Yuhua Zhu
This paper develops a model-based framework for continuous-time policy evaluation (CTPE) in reinforcement learning, incorporating both Brownian and Lévy noise to model stochastic…
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control
Yuhua Zhu, Yutong Ren
In this paper, we develop an off-policy method for continuous-time reinforcement learning (CTRL), where the system dynamics are governed by an unknown stochastic differential equat…
Discretization error from regularized Reinforcement Learning to continuous-time stochastic control
Huyên Pham, Yuming Paul Zhang, Yuhua Zhu
This paper establishes a rigorous connection between regularized discrete-time reinforcement learning (RL) and continuous-time stochastic optimal control. Specifically, classical R…
PhiBE: A PDE-based Bellman Equation for Continuous Time Policy Evaluation
Yuhua Zhu
In this paper, we study policy evaluation in continuous-time reinforcement learning (RL), where the state follows an unknown stochastic differential equation (SDE), but only discre…
Optimal-PhiBE: A PDE-based Model-free framework for Continuous-time Reinforcement Learning
Yuhua Zhu, Yuming Zhang, Haoyu Zhang
This paper addresses continuous-time reinforcement learning (CTRL) where the system dynamics are governed by an unknown stochastic differential equation, and only discrete-time obs…
On Bellman equations for continuous-time policy evaluation I: discretization and approximation
Wenlong Mou, Yuhua Zhu
We study the problem of computing the value function from a discretely-observed trajectory of a continuous-time diffusion process. We develop a new class of algorithms based on eas…