13 citations · 16 across the 4 of their papers we have counts for
4 papers
Multilevel Interpretability Of Artificial Neural Networks: Leveraging Framework And Methods From Neuroscience
Zhonghao He, Jascha Achterberg, Katie Collins +13
As deep learning systems are scaled up to many billions of parameters, relating their internal structure to external behaviors becomes very challenging. Although daunting, this pro…
Universal Neurons in GPT2 Language Models
Wes Gurnee, Theo Horsley, Zifan Carl Guo +5
A basic question within the emerging field of mechanistic interpretability is the degree to which neural networks learn the same underlying mechanisms. In other words, are neural m…
Training Dynamics of Contextual N-Grams in Language Models
Lucia Quirke, Lovis Heindrich, Wes Gurnee +1
Prior work has shown the existence of contextual neurons in language models, including a neuron that activates on German text. We show that this neuron exists within a broader cont…
Finding Neurons in a Haystack: Case Studies with Sparse Probing
Wes Gurnee, Neel Nanda, Matthew Pauly +3
Despite rapid adoption and deployment of large language models (LLMs), the internal computations of these models remain opaque and poorly understood. In this work, we seek to under…