5 papers
TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment
Sweta Mahajan, Sukrut Rao, Jiahao Xie +2
Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and text embeddings are often poorly…
FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
Amin Parchami-Araghi, Sukrut Rao, Jonas Fischer +1
Deep networks have shown remarkable performance across a wide range of tasks, yet getting a global concept-level understanding of how they function remains a key challenge. Many po…
CFM: Language-aligned Concept Foundation Model for Vision
Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi +2
Language-aligned vision foundation models perform strongly across diverse downstream tasks. Yet, their learned representations remain opaque, making interpreting their decision-mak…
B-cos LM: Efficiently Transforming Pre-trained Language Models for Improved Explainability
Yifan Wang, Sukrut Rao, Ji-Ung Lee +2
Post-hoc explanation methods for black-box models often struggle with faithfulness and human interpretability due to the lack of explainability in current neural architectures. Mea…
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
Shreyash Arya, Sukrut Rao, Moritz Böhle +1
B-cos Networks have been shown to be effective for obtaining highly human interpretable explanations of model decisions by architecturally enforcing stronger alignment between inpu…