Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
InceptionNeXt: When Inception Meets ConvNeXt
Weihao Yu, Pan Zhou, Shuicheng Yan +1
Inspired by the long-range modeling ability of ViTs, large-kernel convolutions are widely studied and adopted recently to enlarge the receptive field and improve model performance,…
cs.CV2024
MetaFormer Baselines for Vision
Weihao Yu, Chenyang Si, Pan Zhou +5
MetaFormer, the abstracted architecture of Transformer, has been found to play a significant role in achieving competitive performance. In this paper, we further explore the capaci…
cs.CV2024
MambaOut: Do We Really Need Mamba for Vision?
Weihao Yu, Xinchao Wang
Mamba, an architecture with RNN-like token mixer of state space model (SSM), was recently introduced to address the quadratic complexity of the attention mechanism and subsequently…