1 paper
Fei Deng, Qifei Wang, Wei Wei +2
Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using…