most citedSecrets of RLHF in Large Language Models Part II: Reward Modeling

8 citations · 18 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20246 cited

EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models

Weikang Zhou, Xiao Wang, Limao Xiong +18

Jailbreak attacks are crucial for identifying and mitigating the security vulnerabilities of Large Language Models (LLMs). They are designed to bypass safeguards and elicit prohibi…

cs.CL20241 cited

CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models

Huijie Lv, Xiao Wang, Yuansen Zhang +6

Adversarial misuse, particularly through `jailbreaking' that circumvents a model's safety and ethical protocols, poses a significant challenge for Large Language Models (LLMs). Thi…

cs.SE20243 cited

StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Shihan Dou, Yan Liu, Haoxiang Jia +14

The advancement of large language models (LLMs) has significantly propelled the field of code generation. Previous work integrated reinforcement learning (RL) with compiler feedbac…

cs.CV2024

MouSi: Poly-Visual-Expert Vision-Language Models

Xiaoran Fan, Tao Ji, Changhao Jiang +21

Current large vision-language models (VLMs) often encounter challenges such as insufficient capabilities of a single visual component and excessively long visual tokens. These issu…

cs.AI20248 cited

Secrets of RLHF in Large Language Models Part II: Reward Modeling

Binghai Wang, Rui Zheng, Lu Chen +24

Reinforcement Learning from Human Feedback (RLHF) has become a crucial technology for aligning language models with human values and intentions, enabling models to produce more hel…