1 paper · 1 filter
Sebastian Gerstner, Hinrich Schütze
Interpretability researchers have attempted to understand MLP neurons of language models based on both the contexts in which they activate and their output weight vectors. They hav…