collaborators

5 papers

cs.LG2026

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

Praneet Suresh, Jack Stanley, Sonia Joseph +2

Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face u…

cs.LG2025

From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers

Praneet Suresh, Jack Stanley, Sonia Joseph +2

As generative AI systems become competent and democratized in science, business, and government, deeper insight into their failure modes now poses an acute need. The occasional vol…

cs.CV2025

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Sonia Joseph, Praneet Suresh, Lorenz Hufe +7

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision…

cs.CV2025

Decoding Vision Transformers: the Diffusion Steering Lens

Ryota Takatsuki, Sonia Joseph, Ippei Fujisawa +1

Logit Lens is a widely adopted method for mechanistic interpretability of transformer-based language models, enabling the analysis of how internal representations evolve across lay…

cs.CV2025

Steering CLIP's vision transformer with sparse autoencoders

Sonia Joseph, Praneet Suresh, Ethan Goldfarb +6

While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but whic…