6 papers
Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data
Erhan Bayraktar, Martin Hernandez, Qinxin Yan +1
This paper develops a model-free framework for continuous-time mean-field control when the population evolves according to unknown controlled McKean--Vlasov dynamics and only discr…
Policy Gradient for Continuous-Time Mean-Field Control
Erhan Bayraktar, Martin Hernandez, Qinxin Yan +1
This paper develops a policy gradient method for entropy-regularized mean-field control in the discounted infinite-horizon setting. We consider randomized feedback policies and a c…
Stabilizing Policy Optimization via Logits Convexity
Hongzhan Chen, Tao Yang, Yuhua Zhu +3
While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unstable, especially when compared to super…
Variance Reduction via Resampling and Experience Replay
Jiale Han, Xiaowu Dai, Yuhua Zhu
Experience replay is a foundational technique in reinforcement learning that enhances learning stability by storing past experiences in a replay buffer and reusing them during trai…
On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective
Xiaowu Dai, Yuhua Zhu
We study the statistical properties of the dynamic trajectory of stochastic gradient descent (SGD). We approximate the mini-batch SGD and the momentum SGD as stochastic differentia…
Towards Theoretical Understanding of Large Batch Training in Stochastic Gradient Descent
Xiaowu Dai, Yuhua Zhu
Stochastic gradient descent (SGD) is almost ubiquitously used for training non-convex optimization tasks. Recently, a hypothesis proposed by Keskar et al. [2017] that large batch m…