10 papers
NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics
Qizhen Ying, Guangming Wang, Yangchen Pan +3
Physics-grounded video generation requires controllable 3D object dynamics that remain physically consistent under contact, deformation, and external forcing. Existing trajectory-b…
Temporal Difference Learning for Diffusion Models
Qizhen Ying, Yangchen Pan, Victor Adrian Prisacariu +1
Diffusion models are typically trained with objectives that focus on local denoising targets at individual time steps (or adjacent pairs), which do not enforce consistency between…
A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
Axel Barroso-Laguna, Tommaso Cavallari, Victor Adrian Prisacariu +1
Visually localizing an image, i.e., estimating its camera pose, requires building a scene representation that serves as a visual map. The representation we choose has direct conseq…
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
Xianzheng Ma, Tao Sun, Shuai Chen +7
Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on te…
Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections
Jing Wu, Zirui Wang, Iro Laina +1
Mirror reflections are common in everyday environments and can provide stereo information within a single capture, as the real and reflected virtual views are visible simultaneousl…
When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
Xianzheng Ma, Brandon Smart, Yash Bhalgat +14
As large language models (LLMs) evolve, their integration with 3D spatial data (3D-LLMs) has seen rapid progress, offering unprecedented capabilities for understanding and interact…