2 papers
cs.LG2025
Finding Manifolds With Bilinear Autoencoders
Thomas Dooms, Ward Gauderis
Sparse autoencoders are a standard tool for uncovering interpretable latent representations in neural networks. Yet, their interpretation depends on the inputs, making their isolat…
cs.LG2025
Tokenized SAEs: Disentangling SAE Reconstructions
Thomas Dooms, Daniel Wilhelm
Sparse auto-encoders (SAEs) have become a prevalent tool for interpreting language models' inner workings. However, it is unknown how tightly SAE features correspond to computation…