1 paper
Eric Zhao, Jessica Dai, Pranjal Awasthi
Recent progress in strengthening the capabilities of large language models has stemmed from applying reinforcement learning to domains with automatically verifiable outcomes. A key…