54 citations · 54 across the 1 of their papers we have counts for
1 paper
Kevin Wang, Alexandre Variengien, Arthur Conmy +2
Research in mechanistic interpretability seeks to explain behaviors of machine learning models in terms of their internal components. However, most previous work either focuses on…