Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Scaling Interpretable Transformers with Parity Bottleneck Layers
Andrew Mack, Kraig Yuheng Tou, Mark Henry +2
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are de…
cs.LG2026
Mechanistically Eliciting Latent Behaviors in Language Models
Andrew Mack, Nina Panickssery, Alexander Matt Turner
We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could help reshape model behavior and systemat…