1 paper · 1 filter
Beining Wang, Weihang Su, Hongtao Tian +5
Improving the multi-step reasoning ability of Large Language Models (LLMs) is a critical yet challenging task. The dominant paradigm, outcome-supervised reinforcement learning (RLV…