54 citations · 60 across the 3 of their papers we have counts for
1 paper · 1 filter
Kevin Wang, Alexandre Variengien, Arthur Conmy +2
Research in mechanistic interpretability seeks to explain behaviors of machine learning models in terms of their internal components. However, most previous work either focuses on…