1 paper
Jiacheng Guo, Suozhi Huang, Shuzhen Li +11
Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and co…