activity
20242026
collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

From Mechanistic to Compositional Interpretability

Ward Gauderis, Thomas Dooms, Steven T. Homer +2

Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components. Without a formal fr…

cs.LG2026

When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability

ML Nissen Gonzalez, Melwina Albuquerque, Laurence Wroe +3

Mechanistic interpretability aims to break models into meaningful parts; verifying that two such parts implement the same computation is a prerequisite. Existing similarity measure…

cs.LG2026

Bilinear autoencoders find interpretable manifolds

Thomas Dooms, Ward Gauderis, Geraint Wiggins +1

Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span manifolds that current linea…

cs.LG2025

Finding Manifolds With Bilinear Autoencoders

Thomas Dooms, Ward Gauderis

Sparse autoencoders are a standard tool for uncovering interpretable latent representations in neural networks. Yet, their interpretation depends on the inputs, making their isolat…

cs.LG2025

Bilinear MLPs enable weight-based mechanistic interpretability

Michael T. Pearce, Thomas Dooms, Alice Rigg +2

A mechanistic understanding of how MLPs do computation in deep neural networks remains elusive. Current interpretability work can extract features from hidden activations over an i…

cs.LG2025

Compositionality Unlocks Deep Interpretable Models

Thomas Dooms, Ward Gauderis, Geraint A. Wiggins +1

We propose -net, an intrinsically interpretable architecture combining the compositional multilinear structure of tensor networks with the expressivity and efficiency of deep n…