1 paper
Niloofar Gholipour, Marcos Assuncao, Gursimran Singh +8
Reasoning-oriented reinforcement learning enables large language models to solve mathematical, coding, and other multi-step tasks, but shifts a substantial portion of the training…