1 paper · 1 filter
Wu Fei, Hao Kong, Shuxian Liang +5
Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additiona…