2 citations · 2 across the 6 of their papers we have counts for
4 papers · 1 filter
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
Daniel Wurgaft, Can Rager, Matthew Kowal +13
Neural representations carry rich geometric structure; but does that structure causally shape behavior? To address this question, we intervene along paths through activation space…
Do Sparse Autoencoders Capture Concept Manifolds?
Usha Bhalla, Thomas Fel, Can Rager +9
Sparse autoencoders (SAEs) are widely used to extract interpretable features from neural network representations, often under the implicit assumption that concepts correspond to in…
Structured World Representations in Maze-Solving Transformers
Michael Igorevich Ivanitskiy, Alex F. Spies, Tilman Räuker +9
Transformer models underpin many recent advances in practical machine learning applications, yet understanding their internal behavior continues to elude researchers. Given the siz…
A Configurable Library for Generating and Manipulating Maze Datasets
Michael Igorevich Ivanitskiy, Rusheb Shah, Alex F. Spies +8
Understanding how machine learning models respond to distributional shifts is a key research challenge. Mazes serve as an excellent testbed due to varied generation algorithms offe…