6 papers
Imitating from auxiliary imperfect demonstrations via Adversarial Density Weighted Regression
Ziqi Zhang, Zifeng Zhuang, Jingzehua Xu +4
We propose a novel one-step supervised imitation learning (IL) framework called Adversarial Density Regression (ADR). This IL framework aims to correct the policy learned on unknow…
Nash CoT: Multi-Path Inference with Preference Equilibrium
Ziqi Zhang, Cunxiang Wang, Xiong Xiao +2
Chain of thought (CoT) is a reasoning framework that can enhance the performance of Large Language Models (LLMs) on complex inference tasks. In particular, among various studies re…
A dynamical clipping approach with task feedback for Proximal Policy Optimization
Ziqi Zhang, Jingzehua Xu, Zifeng Zhuang +4
Proximal Policy Optimization (PPO) has been broadly applied to robotics learning, showcasing stable training performance. However, the fixed clipping bound setting may limit the pe…
Reinformer: Max-Return Sequence Modeling for Offline RL
Zifeng Zhuang, Dengyun Peng, Jinxin Liu +2
As a data-driven paradigm, offline reinforcement learning (RL) has been formulated as sequence modeling that conditions on the hindsight information including returns, goal or futu…
Improving Offline-to-Online Reinforcement Learning with Q Conditioned State Entropy Exploration
Ziqi Zhang, Xiao Xiong, Zifeng Zhuang +2
Studying how to fine-tune offline reinforcement learning (RL) pre-trained policy is profoundly significant for enhancing the sample efficiency of RL algorithms. However, directly f…
Context-Former: Stitching via Latent Conditioned Sequence Modeling
Ziqi Zhang, Jingzehua Xu, Jinxin Liu +4
Offline reinforcement learning (RL) algorithms can learn better decision-making compared to behavior policies by stitching the suboptimal trajectories to derive more optimal ones.…