1 paper
Ling Xiao, Yuliang Xiu, Yue Chen +2
A typical 2D-to-3D pipeline takes multi-view images as input, where a Vision Foundation Model (VFM) extracts features that are spatially upsampled to dense representations for 3D r…