1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Beining Wang, Weihang Su, Hongtao Tian +5
Improving the multi-step reasoning ability of Large Language Models (LLMs) is a critical yet challenging task. The dominant paradigm, outcome-supervised reinforcement learning (RLV…