collaborators

5 papers

cs.LG2026

Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

Akshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy +3

Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential r…

cs.CV2025

Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability

Tuomas Oikarinen, Ge Yan, Akshay Kulkarni +1

Interpreting individual neurons or directions in activation space is an important topic in mechanistic interpretability. Numerous automated interpretability methods have been propo…

cs.CL2025

ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability

Chung-En Sun, Ge Yan, Akshay Kulkarni +1

Recent advances in long chain-of-thought (CoT) reasoning have largely prioritized answer accuracy and token efficiency, while overlooking aspects critical to trustworthiness. We ar…

cs.CV2025

Interpretable Generative Models through Post-hoc Concept Bottlenecks

Akshay Kulkarni, Ge Yan, Chung-En Sun +2

Concept bottleneck models (CBM) aim to produce inherently interpretable models that rely on human-understandable concepts for their predictions. However, existing approaches to des…

cs.CV2025

Interpreting Neurons in Deep Vision Networks with Language Models

Nicholas Bai, Rahul A. Iyer, Tuomas Oikarinen +2

In this paper, we propose Describe-and-Dissect (DnD), a novel method to describe the roles of hidden neurons in vision networks. DnD utilizes recent advancements in multimodal deep…