1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Kusha Sareen, Morgane M Moss, Alessandro Sordoni +2
Prevalent reinforcement learning~(RL) methods for fine-tuning LLM reasoners, such as GRPO or Leave-one-out PPO, abandon the learned value function in favor of empirically estimated…