activity
20212023
most citedImproving Vision Transformers for Incremental Learning

5 citations · 23 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CV20234 cited

Learning from Rich Semantics and Coarse Locations for Long-tailed Object Detection

Lingchen Meng, Xiyang Dai, Jianwei Yang +7

Long-tailed object detection (LTOD) aims to handle the extreme data imbalance in real-world datasets, where many tail classes have scarce instances. One popular strategy is to expl…

cs.CV2023

Improving Adversarial Robustness of Masked Autoencoders via Test-time Frequency-domain Prompting

Qidong Huang, Xiaoyi Dong, Dongdong Chen +5

In this paper, we investigate the adversarial robustness of vision transformers that are equipped with BERT pretraining (e.g., BEiT, MAE). A surprising observation is that MAE has…

cs.CV20231 cited

PaintSeg: Training-free Segmentation via Painting

Xiang Li, Chung-Ching Lin, Yinpeng Chen +3

The paper introduces PaintSeg, a new unsupervised method for segmenting objects without any training. We propose an adversarial masked contrastive painting (AMCP) process, which cr…

cs.CV20235 cited

Layer Grafted Pre-training: Bridging Contrastive Learning And Masked Image Modeling For Label-Efficient Representations

Ziyu Jiang, Yinpeng Chen, Mengchen Liu +5

Recently, both Contrastive Learning (CL) and Mask Image Modeling (MIM) demonstrate that self-supervision is powerful to learn good representations. However, naively combining them…

cs.CV20231 cited

Find Beauty in the Rare: Contrastive Composition Feature Clustering for Nontrivial Cropping Box Regression

Zhiyu Pan, Yinpeng Chen, Jiale Zhang +3

Automatic image cropping algorithms aim to recompose images like human-being photographers by generating the cropping boxes with improved composition quality. Cropping box regressi…

cs.CV20224 cited

Video Mobile-Former: Video Recognition with Efficient Global Spatial-temporal Modeling

Rui Wang, Zuxuan Wu, Dongdong Chen +6

Transformer-based models have achieved top performance on major video recognition benchmarks. Benefiting from the self-attention mechanism, these models show stronger ability of mo…