1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2024
Toward Optimal LLM Alignments Using Two-Player Games
Rui Zheng, Hongyi Guo, Zhihan Liu +10
The standard Reinforcement Learning from Human Feedback (RLHF) framework primarily focuses on optimizing the performance of large language models using pre-collected prompts. Howev…
cs.CL2024★ 1 cited
Large Language Models as Agents in Two-Player Games
Yang Liu, Peng Sun, Hang Li
By formally defining the training processes of large language models (LLMs), which usually encompasses pre-training, supervised fine-tuning, and reinforcement learning with human f…