1 paper
Gabriel Franco, Carson Loughridge, Mark Crovella
Identifying feature representations in language models is a central task in mechanistic interpretability. Several recent studies have made the observation that feature representati…