1 paper
Henk Tillman, Dan Mossing
Language models can behave in unexpected and unsafe ways, and so it is valuable to monitor their outputs. Internal activations of language models encode additional information that…