1 paper
Andrew Mack, Kraig Yuheng Tou, Mark Henry +2
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are de…