1 paper
Samyak Jain, Ekdeep Singh Lubana, Kemal Oksuz +4
Safety fine-tuning helps align Large Language Models (LLMs) with human preferences for their safe deployment. To better understand the underlying factors that make models safe via…