2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Jiaqi Zhang, Ashton Lee, Anthony Wong +3
Vision Foundation Models (VFMs) with Vision Transformer (ViT) backbones, such as DINOv2, have become essential for downstream tasks like object recognition and semantic segmentatio…