3 papers
cs.CV2025
Probing Visual Language Priors in VLMs
Tiange Luo, Ang Cao, Gunhee Lee +2
Despite recent advances in Vision-Language Models (VLMs), they may over-rely on visual language priors existing in their training data rather than true visual reasoning. To investi…
cs.CV2024
Meta 3D Gen
Raphael Bensadoun, Tom Monnier, Yanir Kleiman +17
We introduce Meta 3D Gen (3DGen), a new state-of-the-art, fast pipeline for text-to-3D asset generation. 3DGen offers 3D asset creation with high prompt fidelity and high-quality 3…
cs.CV2024
Lightplane: Highly-Scalable Components for Neural 3D Fields
Ang Cao, Justin Johnson, Andrea Vedaldi +1
Contemporary 3D research, particularly in reconstruction and generation, heavily relies on 2D images for inputs or supervision. However, current designs for these 2D-3D mapping are…