5 papers
Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
Huikang Su, Dengyun Peng, Zifeng Zhuang +4
Offline safe reinforcement learning aims to learn policies that satisfy predefined safety constraints from static datasets. Existing sequence-model-based methods condition action g…
Unlock Reliable Skill Inference for Quadruped Adaptive Behavior by Skill Graph
Hongyin Zhang, Diyuan Shi, Zifeng Zhuang +6
Developing robotic intelligent systems that can adapt quickly to unseen wild situations is one of the critical challenges in pursuing autonomous robotics. Although some impressive…
Imitating from auxiliary imperfect demonstrations via Adversarial Density Weighted Regression
Ziqi Zhang, Zifeng Zhuang, Jingzehua Xu +4
We propose a novel one-step supervised imitation learning (IL) framework called Adversarial Density Regression (ADR). This IL framework aims to correct the policy learned on unknow…
Nash CoT: Multi-Path Inference with Preference Equilibrium
Ziqi Zhang, Cunxiang Wang, Xiong Xiao +2
Chain of thought (CoT) is a reasoning framework that can enhance the performance of Large Language Models (LLMs) on complex inference tasks. In particular, among various studies re…
A dynamical clipping approach with task feedback for Proximal Policy Optimization
Ziqi Zhang, Jingzehua Xu, Zifeng Zhuang +4
Proximal Policy Optimization (PPO) has been broadly applied to robotics learning, showcasing stable training performance. However, the fixed clipping bound setting may limit the pe…