1 paper
Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produce…