1 paper
Chase van de Geijn, Ayush Paliwal, Timo Lüddecke +1
Transformers rely on positional encoding to compensate for the inherent permutation invariance of self-attention. Traditional approaches use absolute sinusoidal embeddings or learn…