1 paper
Shruti Joshi, Andrea Dittadi, Sébastien Lachapelle +1
Unsupervised approaches to large language model (LLM) interpretability, such as sparse autoencoders (SAEs), offer a way to decode LLM activations into interpretable and, ideally, c…