1 paper
Zhichao Wang, Andy Wong, Ruslan Belkin
After the pretraining stage of LLMs, techniques such as SFT, RLHF, RLVR, and RFT are applied to enhance instruction-following ability, mitigate undesired responses, improve reasoni…