30 citations · 37 across the 5 of their papers we have counts for
1 paper · 1 filter
Ji Hou, Xiaoliang Dai, Zijian He +2
Current popular backbones in computer vision, such as Vision Transformers (ViT) and ResNets are trained to perceive the world from 2D images. However, to more effectively understan…