Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Interactions Between Crosscoder Features: A Compact Proofs Perspective
Dmitry Manning-Coe, Thomas Read, Anna Soligo +4
Dictionary learning methods like Sparse Autoencoders (SAEs) and crosscoders attempt to explain a model by decomposing its activations into independent features. Interactions betwee…
cs.LG2025
Grokking vs. Learning: Same Features, Different Encodings
Dmitry Manning-Coe, Jacopo Gliozzi, Alexander G. Stapleton +4
Grokking typically achieves similar loss to ordinary, "steady", learning. We ask whether these different learning paths - grokking versus ordinary training - lead to fundamental di…