1 paper
Kevin Zhang, Jingxi Chen, Mohamad Qadri +5
Vision foundation models trained on Internet-scale RGB datasets enable remarkable capabilities across a range of tasks, from text-to-video generation to few-shot 3D scene reconstru…