1 paper · 1 filter
Daniel Zhao, Abhilash Shankarampeta, Lanxiang Hu +2
We propose a novel method that leverages sparse autoencoders (SAEs) and clustering techniques to analyze the internal token representations of large language models (LLMs) and guid…