1 paper · 1 filter
Damiano Fornasiere, Mirko Bronzi, Spencer Kitts +3
We provide evidence that language models can detect, localize and, to a certain degree, verbalize the difference between perturbations applied to their activations. More precisely,…