1 paper
Harrish Thasarathan, Julian Forsyth, Thomas Fel +2
We present Universal Sparse Autoencoders (USAEs), a framework for uncovering and aligning interpretable concepts spanning multiple pretrained deep neural networks. Unlike existing…