5 papers
RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning
Yuexin Bian, Jie Feng, Tao Wang +3
On-policy Reinforcement Learning (RL) remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies,…
Learning Quadruped Walking from Seconds of Demonstration
Ruipeng Zhang, Hongzhan Yu, Ya-Chien Chang +3
Quadruped locomotion provides a natural setting for understanding when model-free learning can outperform model-based control design, by exploiting data patterns to bypass the diff…
When Maximum Entropy Misleads Policy Optimization
Ruipeng Zhang, Ya-Chien Chang, Sicun Gao
The Maximum Entropy Reinforcement Learning (MaxEnt RL) framework is a leading approach for achieving efficient learning and robust performance across many RL tasks. However, MaxEnt…
Improving Value Estimation Critically Enhances Vanilla Policy Gradient
Tao Wang, Ruipeng Zhang, Sicun Gao
Modern policy gradient algorithms, such as TRPO and PPO, outperform vanilla policy gradient in many RL tasks. Questioning the common belief that enforcing approximate trust regions…
SEEV: Synthesis with Efficient Exact Verification for ReLU Neural Barrier Functions
Hongchao Zhang, Zhizhen Qin, Sicun Gao +1
Neural Control Barrier Functions (NCBFs) have shown significant promise in enforcing safety constraints on nonlinear autonomous systems. State-of-the-art exact approaches to verify…