4 papers
Similarity of Neural Network Representations in Superposition
Sunny Liu, Habon Issa, André Longon +4
Comparing internal representations is a central goal in neuroscience and machine learning, but standard linear alignment metrics (Representational Similarity Analysis, Centered Ker…
Adversarial Examples Are Not Bugs, They Are Superposition
Liv Gorton, Owen Lewis
Adversarial examples -- inputs with imperceptible perturbations that fool neural networks -- remain one of deep learning's most perplexing phenomena despite nearly a decade of rese…
Group Crosscoders for Mechanistic Analysis of Symmetry
Liv Gorton
We introduce group crosscoders, an extension of crosscoders that systematically discover and analyse symmetrical features in neural networks. While neural networks often develop eq…
The Missing Curve Detectors of InceptionV1: Applying Sparse Autoencoders to InceptionV1 Early Vision
Liv Gorton
Recent work on sparse autoencoders (SAEs) has shown promise in extracting interpretable features from neural networks and addressing challenges with polysemantic neurons caused by…