374 citations · 581 across the 7 of their papers we have counts for
1 paper · 1 filter
Stephen Zhang, Mustafa Khan, Vardan Papyan
Large language models (LLMs) often concentrate their attention on a few specific tokens referred to as attention sinks. Common examples include the first token, a prompt-independen…