1 paper
Zhenting Wang, Guofeng Cui, Yu-Jhe Li +2
Recent advances in reinforcement learning (RL)-based post-training have led to notable improvements in large language models (LLMs), particularly in enhancing their reasoning capab…