1 paper
Jiaqi Zhang, Ashton Lee, Anthony Wong +3
Vision Foundation Models (VFMs) with Vision Transformer (ViT) backbones, such as DINOv2, have become essential for downstream tasks like object recognition and semantic segmentatio…