1 paper
Yoav Gur-Arieh, Roy Mayan, Chen Agassy +2
Automated interpretability pipelines generate natural language descriptions for the concepts represented by features in large language models (LLMs), such as plants or the first wo…