54 citations · 56 across the 3 of their papers we have counts for
3 papers
Attribution Patching Outperforms Automated Circuit Discovery
Aaquib Syed, Can Rager, Arthur Conmy
Automated interpretability research has recently attracted attention as a potential research direction that could scale explanations of neural network behavior to large models. Exi…
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy +2
Research in mechanistic interpretability seeks to explain behaviors of machine learning models in terms of their internal components. However, most previous work either focuses on…
StyleGAN-induced data-driven regularization for inverse problems
Arthur Conmy, Subhadip Mukherjee, Carola-Bibiane Schönlieb
Recent advances in generative adversarial networks (GANs) have opened up the possibility of generating high-resolution photo-realistic images that were impossible to produce previo…