1 paper · 1 filter
Angie Boggust, Donghao Ren, Yannick Assogba +3
Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions can be vague…