3 papers
cs.LG2026
Scaling Interpretable Transformers with Parity Bottleneck Layers
Andrew Mack, Kraig Yuheng Tou, Mark Henry +2
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are de…
cs.LG2026
Mechanistically Eliciting Latent Behaviors in Language Models
Andrew Mack, Nina Panickssery, Alexander Matt Turner
We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could help reshape model behavior and systemat…
hep-th2026
Towards Worst-Case Guarantees with Scale-Aware Interpretability
Lauren Greenspan, David Berman, Aryeh Brill +9
Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, exp…