1 paper
Ying Peng, Hongsen Ye, Changxin Huang +3
Vision Transformers (ViTs) have achieved strong performance in video action recognition, but their high computational cost limits their practicality. Lightweight CNNs are more effi…