1 paper
Mohammad Beigi, Ming Jin, Junshan Zhang +3
Reinforcement Learning from Human Feedback (RLHF) enables powerful LLM alignment but can introduce reward hacking - models exploit spurious correlations in proxy rewards without ge…