1 paper
Yuanwen Yue, Anurag Das, Francis Engelmann +2
Current visual foundation models are trained purely on unstructured 2D data, limiting their understanding of 3D structure of objects and scenes. In this work, we show that fine-tun…