7 papers
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
Yohan Mathew, Ollie Matthews, Robert McCarthy +4
The rapid proliferation of frontier model agents promises significant societal advances but also raises concerns about systemic risks arising from unsafe interactions. Collusion to…
Punctuation and Predicates in Language Models
Sonakshi Chauhan, Maheep Chaudhary, Koby Choy +2
In this paper we explore where information is collected and how it is propagated throughout layers in large language models (LLMs). We begin by examining the surprising computation…
Training Neural Networks for Modularity aids Interpretability
Satvik Golechha, Dylan Cope, Nandi Schoots
An approach to improve network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We find pretrained models t…
Studying Cross-cluster Modularity in Neural Networks
Satvik Golechha, Maheep Chaudhary, Joan Velja +2
An approach to improve neural network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We define a measure…
Relating Piecewise Linear Kolmogorov Arnold Networks to ReLU Networks
Nandi Schoots, Mattia Jacopo Villani, Niels uit de Bos
Kolmogorov-Arnold Networks are a new family of neural network architectures which holds promise for overcoming the curse of dimensionality and has interpretability benefits (arXiv:…
Open Problems in Mechanistic Interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson +26
Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…