1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Mingyue Cui, Linghui Shen, Xingyi Yang
Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rely on these decompositions, assuming that…