1 paper · 1 filter
Ananth Eswar, Pratinav Seth, Utsav Avaiya +1
Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they i…