5 papers
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
Eric Onyame, Runtao Zhou, Kowshik Thopalli +2
Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. However, its reliability remains lar…
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
James Flora, Kowshik Thopalli, Akshay R. Kulkarni +2
We present LatentDiff, a scalable framework for semantic dataset comparison that operates directly in the latent space of pretrained vision encoders. By combining sparse autoencode…
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
Akshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy +3
Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential r…
ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
Aditya Ranganath, Hasin Us Sami, Kowshik Thopalli +2
Protein language models often take into consideration the alignment between a protein sequence and its textual description. However, they do not take structural information into co…
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
Xinyuan Yan, Shusen Liu, Kowshik Thopalli +1
Sparse autoencoders (SAEs) have emerged as a powerful tool for uncovering interpretable features in large language models (LLMs) through the sparse directions they learn. However,…