29 citations · 38 across the 2 of their papers we have counts for
3 papers
cs.CV2021★ 9 cited
Rethinking Token-Mixing MLP for MLP-based Vision Backbone
Tan Yu, Xu Li, Yunfeng Cai +2
In the past decade, we have witnessed rapid progress in the machine vision backbone. By introducing the inductive bias from the image processing, convolution neural network (CNN) h…
cs.CV2021★ 29 cited
S-MLP: Spatial-Shift MLP Architecture for Vision
Tan Yu, Xu Li, Yunfeng Cai +2
Recently, visual Transformer (ViT) and its following works abandon the convolution and exploit the self-attention operation, attaining a comparable or even higher accuracy than CNN…
cs.CV2020
STH: Spatio-Temporal Hybrid Convolution for Efficient Action Recognition
Xu Li, Jingwen Wang, Lin Ma +4
Effective and Efficient spatio-temporal modeling is essential for action recognition. Existing methods suffer from the trade-off between model performance and model complexity. In…