33 citations · 51 across the 6 of their papers we have counts for
3 papers · 1 filter
Can Vision Transformers Perform Convolution?
Shanda Li, Xiangning Chen, Di He +1
Several recent studies have demonstrated that attention-based networks, such as Vision Transformer (ViT), can outperform Convolutional Neural Networks (CNNs) on several computer vi…
2.5D Visual Relationship Detection
Yu-Chuan Su, Soravit Changpinyo, Xiangning Chen +8
Visual 2.5D perception involves understanding the semantics and geometry of a scene through reasoning about object relationships with respect to the viewer in an environment. Howev…
Robust and Accurate Object Detection via Adversarial Learning
Xiangning Chen, Cihang Xie, Mingxing Tan +3
Data augmentation has become a de facto component for training high-performance deep image classifiers, but its potential is under-explored for object detection. Noting that most s…