8 papers
From Mechanistic to Compositional Interpretability
Ward Gauderis, Thomas Dooms, Steven T. Homer +2
Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components. Without a formal fr…
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
ML Nissen Gonzalez, Melwina Albuquerque, Laurence Wroe +3
Mechanistic interpretability aims to break models into meaningful parts; verifying that two such parts implement the same computation is a prerequisite. Existing similarity measure…
Bilinear autoencoders find interpretable manifolds
Thomas Dooms, Ward Gauderis, Geraint Wiggins +1
Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span manifolds that current linea…
Finding Manifolds With Bilinear Autoencoders
Thomas Dooms, Ward Gauderis
Sparse autoencoders are a standard tool for uncovering interpretable latent representations in neural networks. Yet, their interpretation depends on the inputs, making their isolat…
Bilinear MLPs enable weight-based mechanistic interpretability
Michael T. Pearce, Thomas Dooms, Alice Rigg +2
A mechanistic understanding of how MLPs do computation in deep neural networks remains elusive. Current interpretability work can extract features from hidden activations over an i…
Parameterized Synthetic Text Generation with SimpleStories
Lennart Finke, Chandan Sreedhara, Thomas Dooms +6
We present SimpleStories, a large synthetic story dataset in simple language, consisting of 2 million samples each in English and Japanese. Through parameterizing prompts at multip…