1 paper · 1 filter
Satchel Grant, Simon Jerome Han, Alexa R. Tartaglini +1
A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to understand what those representations encod…