1 paper · 1 filter
Enrico Cassano, Riccardo Renzulli, Rayyan Ahmed +2
Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated features can serve as controllable…