1 paper
Andrew DiGiugno, Ausif Mahmood
Transformer models typically calculate attention matrices using dot products, which have limitations when capturing nonlinear relationships between embedding vectors. We propose Ne…