1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Ouns El Harzli, Hugo Wallner, Yoonsoo Nam +1
Sparse auto-encoders (SAEs) have re-emerged as a prominent method for mechanistic interpretability, yet they face two significant challenges: the non-smoothness of the L1 penalt…