1 paper
Bowen Zhang, Liyang Liu, Minh Hieu Phan +3
This paper investigates the capability of plain Vision Transformers (ViTs) for semantic segmentation using the encoder-decoder framework and introduces \textbf{SegViTv2}. In this s…