1 paper
Antonio BÄrbÄlau, Cristian Daniel PÄduraru, Teodor Poncu +2
Sparse Autoencoders (SAEs) are widely employed for mechanistic interpretability and model steering. Within this context, steering is by design performed by means of decoding altere…