2 citations · 2 across the 3 of their papers we have counts for
1 paper · 2 filters
Jatin Prakash, Anirudh Buvanesh
Reinforcement learning (RL) with outcome-based rewards has proven effective for improving large language models (LLMs) on complex reasoning tasks. However, its success often depend…