1 paper · 1 filter
Federico Barbero, Ãlvaro Arroyo, Xiangming Gu +4
Large Language Models (LLMs) tend to attend heavily to the first token in the sequence -- creating a so-called attention sink. Many works have studied this phenomenon in detail, pr…