1 paper · 1 filter
Zizhuo Fu, Wenxuan Zeng, Runsheng Wang +1
Large Language Models (LLMs) often assign disproportionate attention to the first token, a phenomenon known as the attention sink. Several recent approaches aim to address this iss…