2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Fei Deng, Qifei Wang, Wei Wei +2
Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using…