1 paper
Xu Wang, Bingqing Jiang, Yu Wan +3
Sparse autoencoders (SAEs) have become a standard tool for mechanistic interpretability in autoregressive large language models (LLMs), enabling researchers to extract sparse, huma…