1 paper
Weiduo Liao, Yunqiao Yang, Ying Wei
Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes…