From the 1 of 3 linked papers with an AI index.
1 paper · 1 filter
Fady Rezk, Yuangang Pan, Chuan-Sheng Foo +4
Personalized alignment from preference data has focused primarily on improving personal reward model (RM) accuracy, with the implicit assumption that better preference ranking tran…