4 citations · 6 across the 5 of their papers we have counts for
5 papers
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
Zhihan Liu, Shenao Zhang, Yongfei Liu +3
Direct preference learning offers a promising and computation-efficient beyond supervised fine-tuning (SFT) for improving code generation in coding large language models (LMs). How…
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
Hongyi Guo, Zhihan Liu, Yufeng Zhang +1
Large Language Models (LLMs) harness extensive data from the Internet, storing a broad spectrum of prior knowledge. While LLMs have proven beneficial as decision-making aids, their…
A Principled Framework for Knowledge-enhanced Large Language Model
Saizhuo Wang, Zhihan Liu, Zhaoran Wang +1
Large Language Models (LLMs) are versatile, yet they often falter in tasks requiring deep and reliable reasoning due to issues like hallucinations, limiting their applicability in…
Sample-Efficient Multi-Agent RL: An Optimization Perspective
Nuoya Xiong, Zhihan Liu, Zhaoran Wang +1
We study multi-agent reinforcement learning (MARL) for the general-sum Markov Games (MGs) under the general function approximation. In order to find the minimum assumption for samp…
Guarded Policy Optimization with Imperfect Online Demonstrations
Zhenghai Xue, Zhenghao Peng, Quanyi Li +2
The Teacher-Student Framework (TSF) is a reinforcement learning setting where a teacher agent guards the training of a student agent by intervening and providing online demonstrati…