58 citations · 65 across the 8 of their papers we have counts for
1 paper · 1 filter
Fan Wu, Huseyin A. Inan, Arturs Backurs +3
Positioned between pre-training and user deployment, aligning large language models (LLMs) through reinforcement learning (RL) has emerged as a prevailing strategy for training ins…