works on

From the 1 of 4 linked papers with an AI index.

collaborators
Showing math.OCShow all

7 papers · 1 filter

math.OC2026

Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies

Ziheng Cheng, Xin Guo, Huyên Pham +1

The paper proposes a model‑free reinforcement learning framework for continuous‑time extended mean field control using deterministic feedback policies, deriving deterministic polic…

math.OC2026

Policy Gradient Learning for Distributionally Robust Markov Decision Processes under Wasserstein Ambiguity

Yadh Hafsi, Samy Mekkaoui, Huyên Pham +1

We study finite-horizon Markov decision processes under distributional uncertainty in the transition kernels and develop a policy-gradient framework for Wasserstein distributionall…

math.OC2026

Discretization error from regularized Reinforcement Learning to continuous-time stochastic control

Huyên Pham, Yuming Paul Zhang, Yuhua Zhu

This paper establishes a rigorous connection between regularized discrete-time reinforcement learning (RL) and continuous-time stochastic optimal control. Specifically, classical R…

math.OC2026

Model-free policy gradient for discrete-time mean-field control

Matthieu Meunier, Huyên Pham, Christoph Reisinger

We study model-free policy learning for discrete-time mean-field control (MFC) problems with finite state space and compact action space. In contrast to the extensive literature on…

math.OC2024

Full error analysis of policy gradient learning algorithms for exploratory linear quadratic mean-field control problem in continuous time with common noise

Noufel Frikha, Huyên Pham, Xuanye Song

We consider reinforcement learning (RL) methods for finding optimal policies in linear quadratic (LQ) mean field control (MFC) problems over an infinite horizon in continuous time,…

math.OC2024

Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching

Robert Denkert, Huyên Pham, Xavier Warin

We propose a comprehensive framework for policy gradient methods tailored to continuous time reinforcement learning. This is based on the connection between stochastic control prob…