1 paper
Enrico Cassano, Riccardo Renzulli, Rayyan Ahmed +2
Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated features can serve as controllable…