4 citations · 7 across the 4 of their papers we have counts for
1 paper · 2 filters
Mehul Damani, Isha Puri, Stewart Slocum +4
When language models (LMs) are trained via reinforcement learning (RL) to generate natural language "reasoning chains", their performance improves on a variety of difficult questio…