1 paper · 1 filter
Qian Lin, Daniel S. Brown
Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common…