collaborators

7 papers

cs.LG2026

Targeted Recovery of Weight-Space Mechanisms From Neural Networks

Antoine Vigouroux, Lee Sharkey

Parameter decomposition (PD) decomposes neural networks into interpretable computational components that faithfully reflect the original network's operations. However, scaling PD t…

cs.LG2025

Stochastic Parameter Decomposition

Lucius Bushnaq, Dan Braun, Lee Sharkey

A key step in reverse engineering neural networks is to decompose them into simpler parts that can be studied in relative isolation. Linear parameter decomposition -- a framework t…

cs.CV2025

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Sonia Joseph, Praneet Suresh, Lorenz Hufe +7

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision…

cs.CY2025

AI Behind Closed Doors: a Primer on The Governance of Internal Deployment

Charlotte Stix, Matteo Pistillo, Girish Sastry +6

The most advanced future AI systems will first be deployed inside the frontier AI companies developing them. According to these companies and independent experts, AI systems may re…

cs.LG2025

Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition

Brianna Chrisman, Lucius Bushnaq, Lee Sharkey

Much of mechanistic interpretability has focused on understanding the activation spaces of large neural networks. However, activation space-based approaches reveal little about the…

cs.LG2025

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition

Dan Braun, Lucius Bushnaq, Stefan Heimersheim +2

Mechanistic interpretability aims to understand the internal mechanisms learned by neural networks. Despite recent progress toward this goal, it remains unclear how best to decompo…