1 paper · 1 filter
T. Ed Li, Junyu Ren
Understanding the internal representations of large language models is crucial for ensuring their reliability and safety, with sparse autoencoders (SAEs) emerging as a promising in…