2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.LG2024
Explorations of Self-Repair in Language Models
Cody Rushing, Neel Nanda
Prior interpretability research studying narrow distributions has preliminarily identified self-repair, a phenomena where if components in large language models are ablated, later…
cs.LG2023★ 2 cited
Copy Suppression: Comprehensively Understanding an Attention Head
Callum McDougall, Arthur Conmy, Cody Rushing +2
We present a single attention head in GPT-2 Small that has one main role across the entire training distribution. If components in earlier layers predict a certain token, and this…