1 paper
Daniel J. Lee, Stefan Heimersheim
Sensitive directions experiments attempt to understand the computational features of Language Models (LMs) by measuring how much the next token prediction probabilities change by p…