2 papers
cs.CL2025
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
Mark Muchane, Sean Richardson, Kiho Park +1
Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A bas…
cs.LG2025
Does Editing Provide Evidence for Localization?
Zihao Wang, Victor Veitch
A basic aspiration for interpretability research in large language models is to "localize" semantically meaningful behaviors to particular components within the LLM. There are vari…