4 papers
Endogenous Resistance to Activation Steering in Language Models
Alex McKenzie, Keenan Pepper, Stijn Servaes +6
Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.g., ``wait, that's not right'') and continuing on-t…
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
Keenan Pepper, Alex McKenzie, Florin Pop +6
Self-interpretation methods prompt language models to describe their own internal states, but remain unreliable due to hyperparameter sensitivity. We show that training lightweight…
Testing Components of the Attention Schema Theory in Artificial Neural Networks
Kathryn T. Farrell, Kirsten Ziman, Michael S. A. Graziano
Growing evidence suggests that the brain uses an attention schema, or a simplified model of attention, to help control what it attends to. One proposed benefit of this model is to…
Unexpected Benefits of Self-Modeling in Neural Systems
Vickram N. Premakumar, Michael Vaiana, Florin Pop +4
Self-models have been a topic of great interest for decades in studies of human cognition and more recently in machine learning. Yet what benefits do self-models confer? Here we sh…