1 paper
Almog Hilel, Riddhi Bhagwat, Idan Shenfeld +2
We describe a vulnerability in language models (LMs) trained with user feedback, whereby a single user can persistently alter LM knowledge and behavior given only the ability to pr…