2 citations · 2 across the 1 of their papers we have counts for
1 paper
Callum McDougall, Arthur Conmy, Cody Rushing +2
We present a single attention head in GPT-2 Small that has one main role across the entire training distribution. If components in earlier layers predict a certain token, and this…