1 paper
Tue M. Cao, Nguyen Do, My T. Thai
Sparse autoencoders (SAEs) have become a central tool for interpreting language models. However, two key SAE analyses that remain difficult to scale are (1) matching semantically s…