1 paper
Raito Kiya, Satoki Ohashi, Kosuke Sato +6
Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (MAs) at the initial position of a sequence. These phenomena frequently…