activity
20152025
most citedInstruction Tuning with GPT-4

189 citations · 683 across the 30 of their papers we have counts for

collaborators
Showing 2023Show all

11 papers · 1 filter

cs.CL2023

Teaching Language Models to Self-Improve through Interactive Demonstrations

Xiao Yu, Baolin Peng, Michel Galley +2

The self-improving ability of large language models (LLMs), enabled by prompting them to analyze and revise their own outputs, has garnered significant interest in recent research.…

cs.CL20231 cited

The Trickle-down Impact of Reward (In-)consistency on RLHF

Lingfeng Shen, Sihao Chen, Linfeng Song +5

Standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for…

cs.CL20232 cited

Stabilizing RLHF through Advantage Model and Selective Rehearsal

Baolin Peng, Linfeng Song, Ye Tian +3

Large Language Models (LLMs) have revolutionized natural language processing, yet aligning these models with human values and preferences using RLHF remains a significant challenge…

cs.CL20232 cited

Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection

Zekun Li, Baolin Peng, Pengcheng He +1

Large Language Models (LLMs) have demonstrated exceptional proficiency in instruction-following, becoming increasingly crucial across various applications. However, this capability…

cs.CL20232 cited

SGP-TOD: Building Task Bots Effortlessly via Schema-Guided LLM Prompting

Xiaoying Zhang, Baolin Peng, Kun Li +2

Building end-to-end task bots and maintaining their integration with new functionalities using minimal human efforts is a long-standing challenge in dialog research. Recently large…

cs.CL2023

Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models

Miaoran Li, Baolin Peng, Michel Galley +2

Fact-checking is an essential task in NLP that is commonly utilized for validating the factual accuracy of claims. Prior work has mainly focused on fine-tuning pre-trained language…