1 paper · 1 filter
Arka Pal, Deep Karkhanis, Samuel Dooley +3
Direct Preference Optimisation (DPO) is effective at significantly improving the performance of large language models (LLMs) on downstream tasks such as reasoning, summarisation, a…