1 paper
Rasul Tutnov, Antoine Grosnit, Haitham Bou-Ammar
Post-alignment of large language models (LLMs) is critical in improving their utility, safety, and alignment with human intentions. Direct preference optimisation (DPO) has become…