4 papers
Holder Policy Optimisation
Yuxiang Chen, Dingli Liang, Yihang Chen +8
Group Relative Policy Optimisation (GRPO) enhances large language models by estimating advantages across a group of sampled trajectories. However, mapping these trajectory-level ad…
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
Xingtu Liu, Lin F. Yang, Sharan Vaswani
We consider infinite-horizon -discounted (linear) constrained Markov decision processes (CMDPs) where the objective is to find a policy that maximizes the expected cumulative r…
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
Yukuan Wei, Xudong Li, Lin F. Yang
Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model. H…
Confident Natural Policy Gradient for Local Planning in -realizable Constrained MDPs
Tian Tian, Lin F. Yang, Csaba Szepesvári
The constrained Markov decision process (CMDP) framework emerges as an important reinforcement learning approach for imposing safety or other critical objectives while maximizing c…