most citedNeuron to Graph: Interpreting Language Model Neurons at Scale

3 citations · 4 across the 2 of their papers we have counts for

collaborators

10 papers

cs.CR202439 cited

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Evan Hubinger, Carson Denison, Jesse Mu +36

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when giv…

cs.AI2024

Large Language Models Relearn Removed Concepts

Michelle Lo, Shay B. Cohen, Fazl Barez

Advances in model editing through neuron pruning hold promise for removing undesirable concepts from large language models. However, it remains unclear whether models have the capa…

cs.AI2023

AI Systems of Concern

Kayla Matteucci, Shahar Avin, Fazl Barez +1

Concerns around future dangers from advanced AI often centre on systems hypothesised to have intrinsic characteristics such as agent-like behaviour, strategic awareness, and long-r…

cs.CL20231 cited

Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark

Jason Hoelscher-Obermaier, Julia Persson, Esben Kran +2

Recent model editing techniques promise to mitigate the problem of memorizing false or outdated associations during LLM training. However, we show that these techniques can introdu…

cs.LG20233 cited

Neuron to Graph: Interpreting Language Model Neurons at Scale

Alex Foote, Neel Nanda, Esben Kran +3

Advances in Large Language Models (LLMs) have led to remarkable capabilities, yet their inner mechanisms remain largely unknown. To understand these models, we need to unravel the…

cs.CL20231 cited

The Larger They Are, the Harder They Fail: Language Models do not Recognize Identifier Swaps in Python

Antonio Valerio Miceli-Barone, Fazl Barez, Ioannis Konstas +1

Large Language Models (LLMs) have successfully been applied to code generation tasks, raising the question of how well these models understand programming. Typical programming lang…