3 citations · 4 across the 2 of their papers we have counts for
3 papers
cs.CL2023★ 1 cited
Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
Jason Hoelscher-Obermaier, Julia Persson, Esben Kran +2
Recent model editing techniques promise to mitigate the problem of memorizing false or outdated associations during LLM training. However, we show that these techniques can introdu…
cs.LG2023★ 3 cited
Neuron to Graph: Interpreting Language Model Neurons at Scale
Alex Foote, Neel Nanda, Esben Kran +3
Advances in Large Language Models (LLMs) have led to remarkable capabilities, yet their inner mechanisms remain largely unknown. To understand these models, we need to unravel the…
cs.LG2023
N2G: A Scalable Approach for Quantifying Interpretable Neuron Representations in Large Language Models
Alex Foote, Neel Nanda, Esben Kran +2
Understanding the function of individual neurons within language models is essential for mechanistic interpretability research. We propose , a tool…