1 paper · 1 filter
Sindhuja Chaduvula, Ahmed Y. Radwan, Azib Farooq +2
Preference alignment methods such as RLHF and Direct Preference Optimization (DPO) improve instruction following, but they can also reinforce hallucinations when preference judgmen…