2 papers
cs.LG2026
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
ML Nissen Gonzalez, Melwina Albuquerque, Laurence Wroe +3
Mechanistic interpretability aims to break models into meaningful parts; verifying that two such parts implement the same computation is a prerequisite. Existing similarity measure…
cs.LG2025
Decomposing The Dark Matter of Sparse Autoencoders
Joshua Engels, Logan Riggs, Max Tegmark
Sparse autoencoders (SAEs) are a promising technique for decomposing language model activations into interpretable linear features. However, current SAEs fall short of completely e…