3 citations · 3 across the 6 of their papers we have counts for
1 paper · 1 filter
Xiangwei Wang, Wei Wang, Ken Chen +2
Reinforcement Learning (RL) serves as a potent paradigm for enhancing reasoning capabilities in Large Language Models (LLMs), yet standard outcome-based approaches often suffer fro…