most citedSecrets of RLHF in Large Language Models Part I: PPO

19 citations · 38 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CL2024

Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data

Han Xia, Songyang Gao, Qiming Ge +3

Reinforcement Learning from Human Feedback (RLHF) has proven effective in aligning large language models with human intentions, yet it often relies on complex methodologies like Pr…

cs.CL20246 cited

EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models

Weikang Zhou, Xiao Wang, Limao Xiong +18

Jailbreak attacks are crucial for identifying and mitigating the security vulnerabilities of Large Language Models (LLMs). They are designed to bypass safeguards and elicit prohibi…

cs.CL2024

Navigating the OverKill in Large Language Models

Chenyu Shi, Xiao Wang, Qiming Ge +7

Large language models are meticulously aligned to be both helpful and harmless. However, recent research points to a potential overkill which means models may refuse to answer beni…

cs.AI20248 cited

Secrets of RLHF in Large Language Models Part II: Reward Modeling

Binghai Wang, Rui Zheng, Lu Chen +24

Reinforcement Learning from Human Feedback (RLHF) has become a crucial technology for aligning language models with human values and intentions, enabling models to produce more hel…

cs.CL20231 cited

RealBehavior: A Framework for Faithfully Characterizing Foundation Models' Human-like Behavior Mechanisms

Enyu Zhou, Rui Zheng, Zhiheng Xi +7

Reports of human-like behaviors in foundation models are growing, with psychological theories providing enduring tools to investigate these behaviors. However, current research ten…

cs.CL20234 cited

TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models

Xiao Wang, Yuansen Zhang, Tianze Chen +9

Aligned large language models (LLMs) demonstrate exceptional capabilities in task-solving, following instructions, and ensuring safety. However, the continual learning aspect of th…