1 paper
Aeree Cho, Alexander D. Greenhalgh, Jonathan Bodea +2
Reinforcement learning has emerged as a dominant technique for fine-tuning the behavior of large language models, with policy optimization (PO) algorithms such as GRPO, DAPO, and D…