1 paper
Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty +2
The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation…