1 paper · 1 filter
Min Yang, Huan Gao, Ping Guo +1
Vision Transformer (ViT) has shown high potential in video recognition, owing to its flexible design, adaptable self-attention mechanisms, and the efficacy of masked pre-training.…