5 papers
At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization
Praneet Suresh, Jack Stanley, Sonia Joseph +2
Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face u…
From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
Praneet Suresh, Jack Stanley, Sonia Joseph +2
As generative AI systems become competent and democratized in science, business, and government, deeper insight into their failure modes now poses an acute need. The occasional vol…
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
Sonia Joseph, Praneet Suresh, Lorenz Hufe +7
Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision…
Decoding Vision Transformers: the Diffusion Steering Lens
Ryota Takatsuki, Sonia Joseph, Ippei Fujisawa +1
Logit Lens is a widely adopted method for mechanistic interpretability of transformer-based language models, enabling the analysis of how internal representations evolve across lay…
Steering CLIP's vision transformer with sparse autoencoders
Sonia Joseph, Praneet Suresh, Ethan Goldfarb +6
While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but whic…