1 paper
Angie Boggust, Donghao Ren, Yannick Assogba +3
Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions can be vague…