1 paper
Hui Lu, Albert Ali Salah, Ronald Poppe
Video understanding requires the extraction of rich spatio-temporal representations, which transformer models achieve through self-attention. Unfortunately, self-attention poses a…