1 paper
Yang You, Yixin Li, Congyue Deng +2
Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehensi…