2 papers
cs.LG2026
Scaling Interpretable Transformers with Parity Bottleneck Layers
Andrew Mack, Kraig Yuheng Tou, Mark Henry +2
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are de…
hep-th2026
Towards Worst-Case Guarantees with Scale-Aware Interpretability
Lauren Greenspan, David Berman, Aryeh Brill +9
Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, exp…