189 citations · 683 across the 30 of their papers we have counts for
11 papers · 1 filter
Teaching Language Models to Self-Improve through Interactive Demonstrations
Xiao Yu, Baolin Peng, Michel Galley +2
The self-improving ability of large language models (LLMs), enabled by prompting them to analyze and revise their own outputs, has garnered significant interest in recent research.…
The Trickle-down Impact of Reward (In-)consistency on RLHF
Lingfeng Shen, Sihao Chen, Linfeng Song +5
Standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for…
Stabilizing RLHF through Advantage Model and Selective Rehearsal
Baolin Peng, Linfeng Song, Ye Tian +3
Large Language Models (LLMs) have revolutionized natural language processing, yet aligning these models with human values and preferences using RLHF remains a significant challenge…
Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection
Zekun Li, Baolin Peng, Pengcheng He +1
Large Language Models (LLMs) have demonstrated exceptional proficiency in instruction-following, becoming increasingly crucial across various applications. However, this capability…
SGP-TOD: Building Task Bots Effortlessly via Schema-Guided LLM Prompting
Xiaoying Zhang, Baolin Peng, Kun Li +2
Building end-to-end task bots and maintaining their integration with new functionalities using minimal human efforts is a long-standing challenge in dialog research. Recently large…
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
Miaoran Li, Baolin Peng, Michel Galley +2
Fact-checking is an essential task in NLP that is commonly utilized for validating the factual accuracy of claims. Prior work has mainly focused on fine-tuning pre-trained language…