1 paper
Huaijie Wang, Shibo Hao, Hanze Dong +4
Improving the multi-step reasoning ability of large language models (LLMs) with offline reinforcement learning (RL) is essential for quickly adapting them to complex tasks. While D…