1 paper · 1 filter
Amit LeVi, Elad David, Max Fomin
Interpretability methods aim to reveal the features represented inside large language models (LLMs). Many existing methods begin with labeled examples of a human-defined concept th…