278 citations · 1.1k across the 26 of their papers we have counts for
13 papers · 2 filters
An Image Patch is a Wave: Phase-Aware Vision MLP
Yehui Tang, Kai Han, Jianyuan Guo +4
In the field of computer vision, recent works show that a pure MLP architecture mainly stacked by fully-connected layers can achieve competing performance with CNN and transformer.…
Learning Versatile Convolution Filters for Efficient Visual Recognition
Kai Han, Yunhe Wang, Chang Xu +3
This paper introduces versatile filters to construct efficient convolutional neural networks that are widely used in various visual recognition tasks. Considering the demands of ef…
Hire-MLP: Vision MLP via Hierarchical Rearrangement
Jianyuan Guo, Yehui Tang, Kai Han +5
Previous vision MLPs such as MLP-Mixer and ResMLP accept linearly flattened image patches as input, making them inflexible for different input sizes and hard to capture spatial inf…
Greedy Network Enlarging
Chuanjian Liu, Kai Han, An Xiao +4
Recent studies on deep convolutional neural networks present a simple paradigm of architecture design, i.e., models with more MACs typically achieve better accuracy, such as Effici…
CMT: Convolutional Neural Networks Meet Vision Transformers
Jianyuan Guo, Kai Han, Han Wu +4
Vision transformers have been successfully applied to image recognition tasks due to their ability to capture long-range dependencies within an image. However, there are still gaps…
Learning Efficient Vision Transformers via Fine-Grained Manifold Distillation
Zhiwei Hao, Jianyuan Guo, Ding Jia +5
In the past few years, transformers have achieved promising performances on various computer vision tasks. Unfortunately, the immense inference overhead of most existing vision tra…