1 paper · 1 filter
Florian Eichin, Yupei Du, Philipp Mondorf +3
Post-hoc interpretability methods typically attribute a model's behavior to its components, data, or training trajectory in isolation, and are often tied to a particular level of g…