1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Sarvesh Shashidhar, Ritik, Nachiketa Patil +2
Direct Preference Optimisation (DPO) has emerged as a powerful method for aligning Large Language Models (LLMs) with human preferences, offering a stable and efficient alternative…