1 paper
Amir Saeidi, Shivanshu Verma, Aswin RRV +2
Reinforcement Learning with Human Feedback (RLHF) enhances the alignment of Large Language Models (LLMs). However, its limitations have led to the development of Direct Preference…