1 paper · 1 filter
Hu Wang, Congbo Ma, Ian Reid +1
The advantage function is a central concept in RL that helps reduce variance in policy gradient estimates. For language modeling, Group Relative Policy Optimization (GRPO) was prop…