1 paper
Yihong Chen, Zhouchen Lin, Quanming Yao
Attention sinks and massive activations are recurring and closely related phenomena in Transformer models. Existing explanations have largely focused on the forward pass, yet in pr…