1 paper · 1 filter
Haichuan Wang, Tao Lin, Lingkai Kong +3
Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.…