32 citations · 91 across the 4 of their papers we have counts for
4 papers
MVT: Multi-view Vision Transformer for 3D Object Recognition
Shuo Chen, Tan Yu, Ping Li
Inspired by the great success achieved by CNN in image recognition, view-based methods applied CNNs to model the projected views for 3D object understanding and achieved excellent…
S-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
Tan Yu, Xu Li, Yunfeng Cai +2
Recently, MLP-based vision backbones emerge. MLP-based vision architectures with less inductive bias achieve competitive performance in image recognition compared with CNNs and vis…
Rethinking Token-Mixing MLP for MLP-based Vision Backbone
Tan Yu, Xu Li, Yunfeng Cai +2
In the past decade, we have witnessed rapid progress in the machine vision backbone. By introducing the inductive bias from the image processing, convolution neural network (CNN) h…
S-MLP: Spatial-Shift MLP Architecture for Vision
Tan Yu, Xu Li, Yunfeng Cai +2
Recently, visual Transformer (ViT) and its following works abandon the convolution and exploit the self-attention operation, attaining a comparable or even higher accuracy than CNN…