1 paper
Hao Wu, Wei Liu
Reinforcement learning has been widely applied to enhance the reasoning capabilities of large language models. Extending the inference limits of smaller models has become a promine…