2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Hao Bai, Yi Ma
Neurons in auto-regressive language models like GPT-2 can be interpreted by analyzing their activation patterns. Recent studies have shown that techniques such as dictionary learni…