paper

Ranking-Augmented On-Policy Optimization with Adaptive Advantage-Normalization for Constrained Control

arXiv:2608.15359

Abstract

This paper analyzes the boundedness and feasibility properties of Advantage-Ranked Group Relative Policy Optimization (A-GRPO), a ranking-augmented, critic-free policy gradient method employing a Transformer-encoder actor for fixed-horizon control with terminal constraints. When feasibility is evaluated only at the final step, the resulting sparse feedback destabilizes critic-based advantage estimation and weakens standard Lagrangian approaches. A trajectory-level ranking mechanism that augments group-relative policy updates by reweighting advantages according to constraint satisfaction is formalized, and three results are established: (i) a scale-adaptive per-timestep normalization bounds advantage variance at every timestep independently, (ii) the ranked advantage strictly separates feasible from violating trajectories under a verifiable ranking-weight condition, biasing the policy gradient toward constraint satisfaction, and (iii) the adaptive dual variables remain bounded and exhibit a drift-balance property that acts as a feedback mechanism for feasibility. These results are validated on a 3,605-step series-hybrid powertrain energy management task with a terminal state-of-charge constraint, where A-GRPO achieves 75.4% mean sustained feasibility with return within 3.7% of the dynamic programming optimum, outperforming a Proximal Policy Optimization with Lagrangian penalties (PPO-Lag) baseline (27.4% sustained), and ablation experiments confirm that both the ranking and Lagrangian components are necessary for this performance.

8 pages, 2 figures. Accepted for presentation at the 2026 IEEE Conference on Decision and Control (CDC). (c) 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media. The full copyright notice appears on the first page of the paper

Ranking-Augmented On-Policy Optimization with Adaptive Advantage-Normalization for Constrained Control · wovepaper