1 paper
Doyeon Lee, Eunyi Lyou, Hyunsoo Cho +3
GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to…