4 papers
The Anatomy of Uncertainty in LLMs
Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli +1
Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment. Current approaches, which either provide a single uncerta…
Improving Robustness In Sparse Autoencoders via Masked Regularization
Vivek Narayanaswamy, Kowshik Thopalli, Bhavya Kailkhura +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability to project LLM activations onto sparse latent spaces. However, sparsity alone is an imperfect proxy for i…
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
Akshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy +3
Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential r…
Leveraging Registers in Vision Transformers for Robust Adaptation
Srikar Yellapragada, Kowshik Thopalli, Vivek Narayanaswamy +5
Vision Transformers (ViTs) have shown success across a variety of tasks due to their ability to capture global image representations. Recent studies have identified the existence o…