collaborators

5 papers

cs.LG2025

Boundary-to-Region Supervision for Offline Safe Reinforcement Learning

Huikang Su, Dengyun Peng, Zifeng Zhuang +4

Offline safe reinforcement learning aims to learn policies that satisfy predefined safety constraints from static datasets. Existing sequence-model-based methods condition action g…

cs.RO2025

Unlock Reliable Skill Inference for Quadruped Adaptive Behavior by Skill Graph

Hongyin Zhang, Diyuan Shi, Zifeng Zhuang +6

Developing robotic intelligent systems that can adapt quickly to unseen wild situations is one of the critical challenges in pursuing autonomous robotics. Although some impressive…

cs.LG2025

Imitating from auxiliary imperfect demonstrations via Adversarial Density Weighted Regression

Ziqi Zhang, Zifeng Zhuang, Jingzehua Xu +4

We propose a novel one-step supervised imitation learning (IL) framework called Adversarial Density Regression (ADR). This IL framework aims to correct the policy learned on unknow…

cs.CL2024

Nash CoT: Multi-Path Inference with Preference Equilibrium

Ziqi Zhang, Cunxiang Wang, Xiong Xiao +2

Chain of thought (CoT) is a reasoning framework that can enhance the performance of Large Language Models (LLMs) on complex inference tasks. In particular, among various studies re…

cs.LG2024

A dynamical clipping approach with task feedback for Proximal Policy Optimization

Ziqi Zhang, Jingzehua Xu, Zifeng Zhuang +4

Proximal Policy Optimization (PPO) has been broadly applied to robotics learning, showcasing stable training performance. However, the fixed clipping bound setting may limit the pe…