1 paper
Giang Son Nguyen, Nhi Ngoc-Yen Nguyen, Wray Buntine +1
Sparse autoencoder (SAE) features are increasingly used to explain and steer language-model behavior, but it remains unclear whether a feature found in one language context plays t…