20 citations · 46 across the 11 of their papers we have counts for
1 paper · 1 filter
Shashata Sawmya, Micah Adler, Nir Shavit
This paper studies the emergence of interpretable categorical features within large language models (LLMs), analyzing their behavior across training checkpoints (time), transformer…