1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Anna Goldie, Azalia Mirhoseini, Hao Zhou +2
Reinforcement learning has been shown to improve the performance of large language models. However, traditional approaches like RLHF or RLAIF treat the problem as single-step. As f…