1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Pedro Sandoval-Segura, Xijun Wang, Ashwinee Panda +4
Attention is foundational to large language models (LLMs), enabling different heads to have diverse focus on relevant input tokens. However, learned behaviors like attention sinks,…