1 paper
Yang Zhan, Yunhao Li, Zhang Chao +2
Recent advancements in reinforcement fine-tuning have significantly improved the reasoning ability of large language models (LLMs). In particular, methods such as group relative po…