1 paper
Abhijnan Nath, Changsoo Jung, Ethan Seefried +1
Traditional RLHF-based LLM alignment methods explicitly maximize the expected rewards from a separate reward model. More recent supervised alignment methods like Direct Preference…