1 paper · 1 filter
Saksham Sahai Srivastava, Vaneet Aggarwal
This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Policy Optimization (PPO), Q-Learnin…