1 paper
Hung Nguyen, Chanho Kim, Fuxin Li
Transformers have recently been popular for learning and inference in the spatial-temporal domain. However, their performance relies on storing and applying attention to the featur…