18 citations · 20 across the 5 of their papers we have counts for
5 papers
Crafting Large Language Models for Enhanced Interpretability
Chung-En Sun, Tuomas Oikarinen, Tsui-Wei Weng
We introduce the Concept Bottleneck Large Language Model (CB-LLM), a pioneering approach to creating inherently interpretable Large Language Models (LLMs). Unlike traditional black…
Corrupting Neuron Explanations of Deep Visual Features
Divyansh Srivastava, Tuomas Oikarinen, Tsui-Wei Weng
The inability of DNNs to explain their black-box behavior has led to a recent surge of explainability methods. However, there are growing concerns that these explainability methods…
The Importance of Prompt Tuning for Automated Neuron Explanations
Justin Lee, Tuomas Oikarinen, Arjun Chatha +3
Recent advances have greatly increased the capabilities of large language models (LLMs), but our understanding of the models and their safety has not progressed as fast. In this pa…
Concept-Monitor: Understanding DNN training through individual neurons
Mohammad Ali Khan, Tuomas Oikarinen, Tsui-Wei Weng
In this work, we propose a general framework called Concept-Monitor to help demystify the black-box DNN training processes automatically using a novel unified embedding space and c…
Label-Free Concept Bottleneck Models
Tuomas Oikarinen, Subhro Das, Lam M. Nguyen +1
Concept bottleneck models (CBM) are a popular way of creating more interpretable neural networks by having hidden layer neurons correspond to human-understandable concepts. However…