1 paper
Tristan Gaudreault, Yongyi Mao
Transformers have become the dominant architecture for sequence modeling by using self-attention to enable expressive and highly parallel processing. However, the resulting quadrat…