Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Activation Scaling for Steering and Interpreting Language Models
Niklas Stoehr, Kevin Du, Vésteinn Snæbjarnarson +3
Given the prompt "Rome is in", can we steer a language model to flip its prediction of an incorrect token "France" to a correct token "Italy" by only multiplying a few relevant act…
cs.CL2021
Classifying Dyads for Militarized Conflict Analysis
Niklas Stoehr, Lucas Torroba Hennigen, Samin Ahbab +2
Understanding the origins of militarized conflict is a complex, yet important undertaking. Existing research seeks to build this understanding by considering bi-lateral relationshi…