3 papers
cs.LG2026
Holder Policy Optimisation
Yuxiang Chen, Dingli Liang, Yihang Chen +8
Group Relative Policy Optimisation (GRPO) enhances large language models by estimating advantages across a group of sampled trajectories. However, mapping these trajectory-level ad…
cs.LG2025
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
Xingtu Liu, Lin F. Yang, Sharan Vaswani
We consider infinite-horizon -discounted (linear) constrained Markov decision processes (CMDPs) where the objective is to find a policy that maximizes the expected cumulative r…
cs.LG2025
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
Yukuan Wei, Xudong Li, Lin F. Yang
Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model. H…