2 papers
cs.LG2026
Polymorphism Is Rotation: Operational Mechanistic Interpretability from a Two-Layer Transformer to Pythia-70m
Jordan F. McCann
Independently trained transformers compute the same function in residual-stream bases that differ by a uniform random rotation on . We call this ph…
cs.LG2026
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
Jordan F. McCann
Sparse autoencoders (SAEs) are now standard tools for decomposing language model activations into interpretable features, and automated interpretability pipelines routinely assign…