From the 1 of 9 linked papers with an AI index.
9 papers
Bridging the Gap between Newton-Raphson Method and Regularized Policy Iteration
Zeyang Li, Chuxiong Hu, Yunan Wang +4
The paper shows that regularized policy iteration in reinforcement learning is mathematically equivalent to applying the Newton‑Raphson method to a smoothed Bellman equation, provi…
On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration
Yujie Yang, Zhilong Zheng, Shengbo Eben Li
Ensuring the safety of environmental exploration is a critical problem in reinforcement learning (RL). While limiting exploration to a feasible zone has become widely accepted as a…
The Feasibility Theory of Constrained Reinforcement Learning: A Tutorial Study
Yujie Yang, Zhilong Zheng, Masayoshi Tomizuka +2
Satisfying safety constraints is a priority concern when solving optimal control problems (OCPs). Due to the existence of infeasibility phenomenon, where a constraint-satisfying so…
Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
Jiaming Zhang, Yujie Yang, Haoning Wang +2
Safe reinforcement learning (safe RL) aims to respect safety requirements while optimizing long-term performance. In many practical applications, however, the problem involves an i…
Off-policy Reinforcement Learning with Model-based Exploration Augmentation
Likun Wang, Xiangteng Zhang, Yinuo Wang +5
Exploration is fundamental to reinforcement learning (RL), as it determines how effectively an agent discovers and exploits the underlying structure of its environment to achieve o…
Jump-Start Reinforcement Learning with Self-Evolving Priors for Extreme Monopedal Locomotion
Ziang Zheng, Guojian Zhan, Shiqi Liu +3
Reinforcement learning (RL) has shown great potential in enabling quadruped robots to perform agile locomotion. However, directly training policies to simultaneously handle dual ex…