1 paper · 1 filter
Moritz Miller, Florent Draye, Bernhard Schölkopf +1
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to suppor…