1 paper
Wentong Li, Zhiyuan Qi, Zichen Zhao +2
Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting their effectiveness on tasks req…