16 citations · 24 across the 11 of their papers we have counts for
1 paper · 2 filters
Mengxuan Hu, Vivek V. Datla, Anoop Kumar +4
Recent advances in alignment techniques such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), and Direct Preference Optimization (DPO) have impro…