collaborators

5 papers

cs.LG2026

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

Praneet Suresh, Jack Stanley, Sonia Joseph +2

Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face u…

cs.LG2026

Quantifying LLM Attention-Head Stability: Implications for Circuit Universality

Karan Bali, Jack Stanley, Praneet Suresh +1

In mechanistic interpretability, recent work scrutinizes transformer "circuits" - sparse, mono or multi layer sub computations, that may reflect human understandable functions. Yet…

cs.LG2025

From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers

Praneet Suresh, Jack Stanley, Sonia Joseph +2

As generative AI systems become competent and democratized in science, business, and government, deeper insight into their failure modes now poses an acute need. The occasional vol…

cs.CV2025

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Sonia Joseph, Praneet Suresh, Lorenz Hufe +7

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision…

cs.CV2025

Steering CLIP's vision transformer with sparse autoencoders

Sonia Joseph, Praneet Suresh, Ethan Goldfarb +6

While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but whic…