1 paper
Victoria R. Li, Jenny Kaufmann, Martin Wattenberg +3
Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this a…