2 citations · 2 across the 1 of their papers we have counts for
1 paper
Aaquib Syed, Can Rager, Arthur Conmy
Automated interpretability research has recently attracted attention as a potential research direction that could scale explanations of neural network behavior to large models. Exi…