1 paper
Haoran Sui, Yaoyuan Jia
Vision Transformers (ViTs) are widely believed to require more labeled data than CNNs for industrial dense prediction. Through controlled experiments on four industrial datasets, w…