activity
20242026
collaborators

6 papers

math.NA2026

A Robust Model-Based Approach for Continuous-Time Policy Evaluation with Unknown Lévy Process Dynamics

Qihao Ye, Xiaochuan Tian, Yuhua Zhu

This paper develops a model-based framework for continuous-time policy evaluation (CTPE) in reinforcement learning, incorporating both Brownian and Lévy noise to model stochastic…

math.OC2026

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control

Yuhua Zhu, Yutong Ren

In this paper, we develop an off-policy method for continuous-time reinforcement learning (CTRL), where the system dynamics are governed by an unknown stochastic differential equat…

math.OC2026

Discretization error from regularized Reinforcement Learning to continuous-time stochastic control

Huyên Pham, Yuming Paul Zhang, Yuhua Zhu

This paper establishes a rigorous connection between regularized discrete-time reinforcement learning (RL) and continuous-time stochastic optimal control. Specifically, classical R…

math.OC2026

PhiBE: A PDE-based Bellman Equation for Continuous Time Policy Evaluation

Yuhua Zhu

In this paper, we study policy evaluation in continuous-time reinforcement learning (RL), where the state follows an unknown stochastic differential equation (SDE), but only discre…

math.OC2025

Optimal-PhiBE: A PDE-based Model-free framework for Continuous-time Reinforcement Learning

Yuhua Zhu, Yuming Zhang, Haoyu Zhang

This paper addresses continuous-time reinforcement learning (CTRL) where the system dynamics are governed by an unknown stochastic differential equation, and only discrete-time obs…

cs.LG2024

On Bellman equations for continuous-time policy evaluation I: discretization and approximation

Wenlong Mou, Yuhua Zhu

We study the problem of computing the value function from a discretely-observed trajectory of a continuous-time diffusion process. We develop a new class of algorithms based on eas…