1 paper
Zhibin Duan, Guowei Rong, Zhuo Li +3
Reward models learned from human preferences are central to aligning large language models (LLMs) via reinforcement learning from human feedback, yet they are often vulnerable to r…