1 paper
Angelica Chen, Sadhika Malladi, Lily H. Zhang +4
Preference learning algorithms (e.g., RLHF and DPO) are frequently used to steer LLMs to produce generations that are more preferred by humans, but our understanding of their inner…