18 citations · 21 across the 10 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024★ 2 cited
Crafting Large Language Models for Enhanced Interpretability
Chung-En Sun, Tuomas Oikarinen, Tsui-Wei Weng
We introduce the Concept Bottleneck Large Language Model (CB-LLM), a pioneering approach to creating inherently interpretable Large Language Models (LLMs). Unlike traditional black…
cs.CL2023
The Importance of Prompt Tuning for Automated Neuron Explanations
Justin Lee, Tuomas Oikarinen, Arjun Chatha +3
Recent advances have greatly increased the capabilities of large language models (LLMs), but our understanding of the models and their safety has not progressed as fast. In this pa…