activity
20182026
collaborators

6 papers

math.OC2026

Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

Erhan Bayraktar, Martin Hernandez, Qinxin Yan +1

This paper develops a model-free framework for continuous-time mean-field control when the population evolves according to unknown controlled McKean--Vlasov dynamics and only discr…

math.OC2026

Policy Gradient for Continuous-Time Mean-Field Control

Erhan Bayraktar, Martin Hernandez, Qinxin Yan +1

This paper develops a policy gradient method for entropy-regularized mean-field control in the discounted infinite-horizon setting. We consider randomized feedback policies and a c…

cs.LG2026

Stabilizing Policy Optimization via Logits Convexity

Hongzhan Chen, Tao Yang, Yuhua Zhu +3

While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unstable, especially when compared to super…

stat.ML2025

Variance Reduction via Resampling and Experience Replay

Jiale Han, Xiaowu Dai, Yuhua Zhu

Experience replay is a foundational technique in reinforcement learning that enhances learning stability by storing past experiences in a replay buffer and reusing them during trai…

cs.LG2021

On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective

Xiaowu Dai, Yuhua Zhu

We study the statistical properties of the dynamic trajectory of stochastic gradient descent (SGD). We approximate the mini-batch SGD and the momentum SGD as stochastic differentia…

stat.ML2018

Towards Theoretical Understanding of Large Batch Training in Stochastic Gradient Descent

Xiaowu Dai, Yuhua Zhu

Stochastic gradient descent (SGD) is almost ubiquitously used for training non-convex optimization tasks. Recently, a hypothesis proposed by Keskar et al. [2017] that large batch m…