1 paper
Saksham Sahai Srivastava, Vaneet Aggarwal
This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Policy Optimization (PPO), Q-Learnin…