2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Ryan Koo, Ian Yang, Vipul Raheja +3
Current reinforcement learning from human feedback (RLHF) pipelines for large language model (LLM) alignment typically assign scalar rewards to sequences, using the final token as…