1 paper
Shangwen Sun, Alfredo Canziani, Yann LeCun +1
We study two recurring phenomena in Transformer language models: massive activations, in which a small number of tokens exhibit extreme outliers in a few channels, and attention si…