1 paper · 1 filter
Dongsheng Wang, Jinsen Zhang, Dawei Su +1
Recently, sparse autoencoders (SAEs) have emerged as a promising technique for interpreting activations in foundation models by disentangling features into a sparse set of concepts…