1 paper · 1 filter
Anna Goldie, Azalia Mirhoseini, Hao Zhou +2
Reinforcement learning has been shown to improve the performance of large language models. However, traditional approaches like RLHF or RLAIF treat the problem as single-step. As f…