1 paper
Dmitry Manning-Coe, Thomas Read, Anna Soligo +4
Dictionary learning methods like Sparse Autoencoders (SAEs) and crosscoders attempt to explain a model by decomposing its activations into independent features. Interactions betwee…