1 paper
Taras Kutsyk, Bartosz ZieliÅski
Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. W…