1 paper
Karan Bali, Jack Stanley, Praneet Suresh +1
In mechanistic interpretability, recent work scrutinizes transformer "circuits" - sparse, mono or multi layer sub computations, that may reflect human understandable functions. Yet…