3 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.LG2023★ 2 cited
Copy Suppression: Comprehensively Understanding an Attention Head
Callum McDougall, Arthur Conmy, Cody Rushing +2
We present a single attention head in GPT-2 Small that has one main role across the entire training distribution. If components in earlier layers predict a certain token, and this…
cs.LG2023★ 3 cited
Neuron to Graph: Interpreting Language Model Neurons at Scale
Alex Foote, Neel Nanda, Esben Kran +3
Advances in Large Language Models (LLMs) have led to remarkable capabilities, yet their inner mechanisms remain largely unknown. To understand these models, we need to unravel the…
cs.LG2023
N2G: A Scalable Approach for Quantifying Interpretable Neuron Representations in Large Language Models
Alex Foote, Neel Nanda, Esben Kran +2
Understanding the function of individual neurons within language models is essential for mechanistic interpretability research. We propose , a tool…