4 papers
From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
Praneet Suresh, Jack Stanley, Sonia Joseph +2
As generative AI systems become competent and democratized in science, business, and government, deeper insight into their failure modes now poses an acute need. The occasional vol…
Decoding Vision Transformers: the Diffusion Steering Lens
Ryota Takatsuki, Sonia Joseph, Ippei Fujisawa +1
Logit Lens is a widely adopted method for mechanistic interpretability of transformer-based language models, enabling the analysis of how internal representations evolve across lay…
Steering CLIP's vision transformer with sparse autoencoders
Sonia Joseph, Praneet Suresh, Ethan Goldfarb +6
While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but whic…
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
Sonia Joseph, Praneet Suresh, Lorenz Hufe +7
Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision…