32 citations · 106 across the 11 of their papers we have counts for
5 papers
MVT: Multi-view Vision Transformer for 3D Object Recognition
Shuo Chen, Tan Yu, Ping Li
Inspired by the great success achieved by CNN in image recognition, view-based methods applied CNNs to model the projected views for 3D object understanding and achieved excellent…
S-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
Tan Yu, Xu Li, Yunfeng Cai +2
Recently, MLP-based vision backbones emerge. MLP-based vision architectures with less inductive bias achieve competitive performance in image recognition compared with CNNs and vis…
Rethinking Token-Mixing MLP for MLP-based Vision Backbone
Tan Yu, Xu Li, Yunfeng Cai +2
In the past decade, we have witnessed rapid progress in the machine vision backbone. By introducing the inductive bias from the image processing, convolution neural network (CNN) h…
S-MLP: Spatial-Shift MLP Architecture for Vision
Tan Yu, Xu Li, Yunfeng Cai +2
Recently, visual Transformer (ViT) and its following works abandon the convolution and exploit the self-attention operation, attaining a comparable or even higher accuracy than CNN…
Learning a Robust Representation via a Deep Network on Symmetric Positive Definite Manifolds
Zhi Gao, Yuwei Wu, Xingyuan Bu +1
Recent studies have shown that aggregating convolutional features of a pre-trained Convolutional Neural Network (CNN) can obtain impressive performance for a variety of visual task…