1 paper · 1 filter
Shashata Sawmya, Micah Adler, Nir Shavit
This paper studies the emergence of interpretable categorical features within large language models (LLMs), analyzing their behavior across training checkpoints (time), transformer…