collaborators

10 papers

cs.CV2026

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics

Qizhen Ying, Guangming Wang, Yangchen Pan +3

Physics-grounded video generation requires controllable 3D object dynamics that remain physically consistent under contact, deformation, and external forcing. Existing trajectory-b…

cs.LG2026

Temporal Difference Learning for Diffusion Models

Qizhen Ying, Yangchen Pan, Victor Adrian Prisacariu +1

Diffusion models are typically trained with objectives that focus on local denoising targets at individual time steps (or adjacent pairs), which do not enforce consistency between…

cs.CV2026

A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features

Axel Barroso-Laguna, Tommaso Cavallari, Victor Adrian Prisacariu +1

Visually localizing an image, i.e., estimating its camera pose, requires building a scene representation that serves as a visual map. The representation we choose has direct conseq…

cs.CL2026

Do 3D Large Language Models Really Understand 3D Spatial Relationships?

Xianzheng Ma, Tao Sun, Shuai Chen +7

Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on te…

cs.CV2026

Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections

Jing Wu, Zirui Wang, Iro Laina +1

Mirror reflections are common in everyday environments and can provide stereo information within a single capture, as the real and reflected virtual views are visible simultaneousl…

cs.CV2025

When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models

Xianzheng Ma, Brandon Smart, Yash Bhalgat +14

As large language models (LLMs) evolve, their integration with 3D spatial data (3D-LLMs) has seen rapid progress, offering unprecedented capabilities for understanding and interact…