1 paper
Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant a…