papers

Publications (17)

cs.CV2025

From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs

Ang Cao, Sergio Arnaud, Oleksandr Maksymets +12

3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes--a six-orde…

cs.CV2024

Meta 3D Gen

Raphael Bensadoun, Tom Monnier, Yanir Kleiman +17

We introduce Meta 3D Gen (3DGen), a new state-of-the-art, fast pipeline for text-to-3D asset generation. 3DGen offers 3D asset creation with high prompt fidelity and high-quality 3…

cs.CV2024

Lightplane: Highly-Scalable Components for Neural 3D Fields

Ang Cao, Justin Johnson, Andrea Vedaldi +1

Contemporary 3D research, particularly in reconstruction and generation, heavily relies on 2D images for inputs or supervision. However, current designs for these 2D-3D mapping are…

cs.CV2021

Inverting and Understanding Object Detectors

Ang Cao, Justin Johnson

As a core problem in computer vision, the performance of object detection has improved drastically in the past few years. Despite their impressive performance, object detectors suf…

cs.CV2024

EucliDreamer: Fast and High-Quality Texturing for 3D Models with Stable Diffusion Depth

Cindy Le, Congrui Hetang, Chendi Lin +2

This paper presents a novel method to generate textures for 3D models given text prompts and 3D meshes. Additional depth information is taken into account to perform the Score Dist…

cs.CV2023

Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models

Lukas Höllein, Ang Cao, Andrew Owens +2

We present Text2Room, a method for generating room-scale textured 3D meshes from a given text prompt as input. To this end, we leverage pre-trained 2D text-to-image models to synth…

cs.CV2025

Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

Jianing Yang, Alexander Sax, Kevin J. Liang +6

Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives.…

cs.CV2026

MoE3D: A Mixture-of-Experts Module for 3D Reconstruction

Zichen Wang, Ang Cao, Liam J. Wang +1

We propose a simple yet effective approach to enhance the performance of feed-forward 3D reconstruction models. Existing methods often struggle near depth discontinuities, where st…

cs.CV2024

EucliDreamer: Fast and High-Quality Texturing for 3D Models with Depth-Conditioned Stable Diffusion

Cindy Le, Congrui Hetang, Chendi Lin +2

We present EucliDreamer, a simple and effective method to generate textures for 3D models given text prompts and meshes. The texture is parametrized as an implicit function on the…

cs.CV2025

Probing Visual Language Priors in VLMs

Tiange Luo, Ang Cao, Gunhee Lee +2

Despite recent advances in Vision-Language Models (VLMs), they may over-rely on visual language priors existing in their training data rather than true visual reasoning. To investi…

cs.CV2024

DreamGaussian4D: Generative 4D Gaussian Splatting

Jiawei Ren, Liang Pan, Jiaxiang Tang +4

4D content generation has achieved remarkable progress recently. However, existing methods suffer from long optimization times, a lack of motion controllability, and a low quality…

cs.CV2023

HexPlane: A Fast Representation for Dynamic Scenes

Ang Cao, Justin Johnson

Modeling and re-rendering dynamic 3D scenes is a challenging task in 3D vision. Prior approaches build on NeRF and rely on implicit representations. This is slow since it requires…

eess.SP2019

Unified Signal Compression Using Generative Adversarial Networks

Bowen Liu, Ang Cao, Hun-seok Kim

We propose a unified compression framework that uses generative adversarial networks (GAN) to compress image and speech signals. The compressed signal is represented by a latent ve…

cond-mat.mtrl-sci2026

Building a physics-aware AI ecosystem for solid-state hydrogen storage materials

Seong-Hoon Jang, Yiwen Yao, Chuanyu Liu +66

Hydrogen storage remains a central bottleneck for scalable hydrogen energy systems due to the multiscale and coupled nature of the thermodynamics, kinetics, and microstructural evo…

cs.CV2025

Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D

Sergio Arnaud, Paul McVay, Ada Martin +19

We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state…

cs.CV2022

FWD: Real-time Novel View Synthesis with Forward Warping and Depth

Ang Cao, Chris Rockwell, Justin Johnson

Novel view synthesis (NVS) is a challenging task requiring systems to generate photorealistic images of scenes from new viewpoints, where both quality and speed are important for a…

eess.SP2021

Unified Signal Compression Using a GAN with Iterative Latent Representation Optimization

Bowen Liu, Changwoo Lee, Ang Cao +1

We propose a unified signal compression framework that uses a generative adversarial network (GAN) to compress heterogeneous signals. The compressed signal is represented as a late…