2 citations · 5 across the 7 of their papers we have counts for
1 paper · 1 filter
Siddarth Venkatraman, Matthieu Dinot, Laurence Aitchison
Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their variance reduction strategy. Group-relative methods like GRPO reduce gradient…