32 citations · 106 across the 11 of their papers we have counts for
8 papers · 1 filter
Degenerate Swin to Win: Plain Window-based Transformer without Sophisticated Operations
Tan Yu, Ping Li
The formidable accomplishment of Transformers in natural language processing has motivated the researchers in the computer vision community to build Vision Transformers. Compared w…
R2-MLP: Round-Roll MLP for Multi-View 3D Object Recognition
Shuo Chen, Tan Yu, Ping Li
Recently, vision architectures based exclusively on multi-layer perceptrons (MLPs) have gained much attention in the computer vision community. MLP-like models achieve competitive…
Tree-based Text-Vision BERT for Video Search in Baidu Video Advertising
Tan Yu, Jie Liu, Yi Yang +3
The advancement of the communication technology and the popularity of the smart phones foster the booming of video ads. Baidu, as one of the leading search engine companies in the…
MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting
Xiaoguang Li, Qing Guo, Di Lin +3
Although achieving significant progress, existing deep generative inpainting methods are far from real-world applications due to the low generalization across different scenes. As…
MVT: Multi-view Vision Transformer for 3D Object Recognition
Shuo Chen, Tan Yu, Ping Li
Inspired by the great success achieved by CNN in image recognition, view-based methods applied CNNs to model the projected views for 3D object understanding and achieved excellent…
S-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
Tan Yu, Xu Li, Yunfeng Cai +2
Recently, MLP-based vision backbones emerge. MLP-based vision architectures with less inductive bias achieve competitive performance in image recognition compared with CNNs and vis…