1 paper
Matthew Y. R. Yang, Hao Bai, Ian Wu +3
Outcome-reward reinforcement learning (RL) has proven effective at improving the reasoning capabilities of large language models (LLMs). However, standard RL assigns credit only at…