95 citations · 132 across the 3 of their papers we have counts for
4 papers
Advancing Vision Transformers with Group-Mix Attention
Chongjian Ge, Xiaohan Ding, Zhan Tong +4
Vision Transformers (ViTs) have been shown to enhance visual recognition through modeling long-range dependencies with multi-head self-attention (MHSA), which is typically formulat…
Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
Youwei Liang, Chongjian Ge, Zhan Tong +3
Vision Transformers (ViTs) take all the image patches as tokens and construct multi-head self-attention (MHSA) among them. Complete leverage of these image tokens brings redundant…
MGSampler: An Explainable Sampling Strategy for Video Action Recognition
Yuan Zhi, Zhan Tong, Limin Wang +1
Frame sampling is a fundamental problem in video action recognition due to the essential redundancy in time and limited computation resources. The existing sampling strategy often…
TDN: Temporal Difference Networks for Efficient Action Recognition
Limin Wang, Zhan Tong, Bin Ji +1
Temporal modeling still remains challenging for action recognition in videos. To mitigate this issue, this paper presents a new video architecture, termed as Temporal Difference Ne…