From the 1 of 6 linked papers with an AI index.
6 papers
Bridging the Gap between Newton-Raphson Method and Regularized Policy Iteration
Zeyang Li, Chuxiong Hu, Yunan Wang +4
The paper shows that regularized policy iteration in reinforcement learning is mathematically equivalent to applying the Newton‑Raphson method to a smoothed Bellman equation, provi…
Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning
Yunan Wang, Minghui Song, Zihan Zhang +6
Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL fra…
Time-Optimal Switching Surfaces for Triple Integrator under Full Box Constraints
Yunan Wang, Chuxiong Hu, Zhao Jin
Time-optimal control for triple integrator under full box constraints is a fundamental problem in the field of optimal control, which has been widely applied in the industry. Howev…
Reachability-Augmented Dual Dynamic Programming for Optimal Path Parameterization
Yunan Wang, Jizhou Yan, Chuxiong Hu +1
Optimal path parameterization (OPP) is a fundamental problem for planning trajectories along a prescribed geometric path under kinodynamic constraints and task-dependent objectives…
A Novel State-Centric Necessary Condition for Time-Optimal Control of Controllable Linear Systems Based on Augmented Switching Laws (Extended Version)
Yunan Wang, Chuxiong Hu, Yujie Lin +3
Most existing necessary conditions for optimal control based on adjoining methods require both state and costate information, yet the unobservability of costates for a given feasib…
Chattering Phenomena in Time-Optimal Control for High-Order Chain-of-Integrator Systems with Full State Constraints (Extended Version)
Yunan Wang, Chuxiong Hu, Zeyang Li +3
Time-optimal control for high-order chain-of-integrator systems with full state constraints remains an open and challenging problem within the discipline of optimal control. The be…