138 citations · 142 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 4 cited
Video Mobile-Former: Video Recognition with Efficient Global Spatial-temporal Modeling
Rui Wang, Zuxuan Wu, Dongdong Chen +6
Transformer-based models have achieved top performance on major video recognition benchmarks. Benefiting from the self-attention mechanism, these models show stronger ability of mo…
cs.CV2022★ 138 cited
Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios
Jiashi Li, Xin Xia, Wei Li +6
Due to the complex attention mechanisms and model design, most existing vision Transformers (ViTs) can not perform as efficiently as convolutional neural networks (CNNs) in realist…
cs.CV2021
Gram-SLD: Automatic Self-labeling and Detection for Instance Objects
Rui Wang, Chengtun Wu, Jiawen Xin +1
Instance object detection plays an important role in intelligent monitoring, visual navigation, human-computer interaction, intelligent services and other fields. Inspired by the g…