1 paper · 1 filter
Damek Davis, Benjamin Recht
We show that several popular algorithms for reinforcement learning in large language models with binary rewards can be viewed as stochastic gradient ascent on a monotone transform…