1 paper
Sheng Cao, Mingrui Wu, Karthik Prasad +2
The post-training phase of large language models is essential for enhancing capabilities such as instruction-following, reasoning, and alignment with human preferences. However, it…