From the 1 of 2 linked papers with an AI index.
1 paper · 1 filter
Michael Blum, Mark Silberstein, Yaniv David
Mechanistic interpretability (MI) has emerged as a powerful approach for analyzing and intervening in inference computations, with a growing number of applications such as jailbrea…