1 paper · 1 filter
Niklas Stoehr, Kevin Du, Vésteinn Snæbjarnarson +3
Given the prompt "Rome is in", can we steer a language model to flip its prediction of an incorrect token "France" to a correct token "Italy" by only multiplying a few relevant act…