most citedGuarded Policy Optimization with Imperfect Online Demonstrations

4 citations · 6 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SE2024

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs

Zhihan Liu, Shenao Zhang, Yongfei Liu +3

Direct preference learning offers a promising and computation-efficient beyond supervised fine-tuning (SFT) for improving code generation in coding large language models (LMs). How…

cs.AI20241 cited

Can Large Language Models Play Games? A Case Study of A Self-Play Approach

Hongyi Guo, Zhihan Liu, Yufeng Zhang +1

Large Language Models (LLMs) harness extensive data from the Internet, storing a broad spectrum of prior knowledge. While LLMs have proven beneficial as decision-making aids, their…

cs.CL20231 cited

A Principled Framework for Knowledge-enhanced Large Language Model

Saizhuo Wang, Zhihan Liu, Zhaoran Wang +1

Large Language Models (LLMs) are versatile, yet they often falter in tasks requiring deep and reliable reasoning due to issues like hallucinations, limiting their applicability in…

cs.LG2023

Sample-Efficient Multi-Agent RL: An Optimization Perspective

Nuoya Xiong, Zhihan Liu, Zhaoran Wang +1

We study multi-agent reinforcement learning (MARL) for the general-sum Markov Games (MGs) under the general function approximation. In order to find the minimum assumption for samp…

cs.LG20234 cited

Guarded Policy Optimization with Imperfect Online Demonstrations

Zhenghai Xue, Zhenghao Peng, Quanyi Li +2

The Teacher-Student Framework (TSF) is a reinforcement learning setting where a teacher agent guards the training of a student agent by intervening and providing online demonstrati…