1 paper
Guy Schacht, Ziyad Sheebaelhamd, Riccardo De Santi +2
Preference learning from human feedback has the ability to align generative models with the needs of end-users. Human feedback is costly and time-consuming to obtain, which creates…