4 citations · 8 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 4 cited
Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object Detection
Yuxin Fang, Shusheng Yang, Shijie Wang +3
We present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our t…
cs.CV2020★ 4 cited
Graph Edit Distance Reward: Learning to Edit Scene Graph
Lichang Chen, Guosheng Lin, Shijie Wang +1
Scene Graph, as a vital tool to bridge the gap between language domain and image domain, has been widely adopted in the cross-modality task like VQA. In this paper, we propose a ne…