5 citations · 9 across the 8 of their papers we have counts for
1 paper · 1 filter
Saket Reddy, Andy Liu
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To add…