22 citations · 30 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 4 cited
Learning from Rich Semantics and Coarse Locations for Long-tailed Object Detection
Lingchen Meng, Xiyang Dai, Jianwei Yang +7
Long-tailed object detection (LTOD) aims to handle the extreme data imbalance in real-world datasets, where many tail classes have scarce instances. One popular strategy is to expl…
cs.CV2022★ 4 cited
Video Mobile-Former: Video Recognition with Efficient Global Spatial-temporal Modeling
Rui Wang, Zuxuan Wu, Dongdong Chen +6
Transformer-based models have achieved top performance on major video recognition benchmarks. Benefiting from the self-attention mechanism, these models show stronger ability of mo…
cs.CV2022★ 22 cited
TinyViT: Fast Pretraining Distillation for Small Vision Transformers
Kan Wu, Jinnian Zhang, Houwen Peng +4
Vision transformer (ViT) recently has drawn great attention in computer vision due to its remarkable model capability. However, most prevailing ViT models suffer from huge number o…