1 paper
Haoqi Wang, Tong Zhang, Mathieu Salzmann
Large transformer models are known to produce high-norm tokens. In vision transformers (ViTs), such tokens have been mathematically modeled through the singular vectors of the line…