1 paper
Mohammed Abu Baker, Luca Baroni, Dan Wilhelm
Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors. To study these risks, researchers develop model organi…