40 citations · 64 across the 14 of their papers we have counts for
1 paper · 1 filter
Shashata Sawmya, Micah Adler, Nir Shavit
This paper studies the emergence of interpretable categorical features within large language models (LLMs), analyzing their behavior across training checkpoints (time), transformer…