5 papers
Improving Robustness In Sparse Autoencoders via Masked Regularization
Vivek Narayanaswamy, Kowshik Thopalli, Bhavya Kailkhura +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability to project LLM activations onto sparse latent spaces. However, sparsity alone is an imperfect proxy for i…
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
Akshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy +3
Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential r…
ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
Aditya Ranganath, Hasin Us Sami, Kowshik Thopalli +2
Protein language models often take into consideration the alignment between a protein sequence and its textual description. However, they do not take structural information into co…
VERIRAG: A Post-Retrieval Auditing of Scientific Study Summaries
Shubham Mohole, Hongjun Choi, Shusen Liu +6
Can democratized information gatekeepers and community note writers effectively decide what scientific information to amplify? Lacking domain expertise, such gatekeepers rely on au…
Leveraging Registers in Vision Transformers for Robust Adaptation
Srikar Yellapragada, Kowshik Thopalli, Vivek Narayanaswamy +5
Vision Transformers (ViTs) have shown success across a variety of tasks due to their ability to capture global image representations. Recent studies have identified the existence o…