most citedCooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models

1 citations · 2 across the 20 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter