1 paper · 1 filter
Wei Shi, Sihang Li, Tao Liang +4
Mechanistic interpretability of large language models (LLMs) aims to uncover the internal processes of information propagation and reasoning. Sparse autoencoders (SAEs) have demons…