14 citations · 14 across the 3 of their papers we have counts for
1 paper · 1 filter
T. Ed Li, Junyu Ren
Understanding the internal representations of large language models is crucial for ensuring their reliability and safety, with sparse autoencoders (SAEs) emerging as a promising in…