5 papers
ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures
Haoran Xu, Lechao Zhang, Daoguo Dong +2
Constructing simulation-ready 3D scenes from multi-view captures is a key bottleneck for Embodied Artificial Intelligence, as downstream tasks require object-level structure, expli…
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
Zizhao Tong, Yeying Jin, Hongfeng Lai +11
Interactive world models for first-person shooter (FPS) games must resolve high-frequency overlapping control signals at every frame without disrupting unaffected regions. Existing…
PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning
Dongxu Zhang, Yiding Sun, Pengcheng Li +12
While Multimodal Large Language Models (MLLMs) demonstrate proficiency in 2D scenes, extending their perceptual intelligence to 3D point cloud understanding remains a significant c…
Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing
Weitao Wang, Haoran Xu, Jun Meng +1
As 3D generation techniques continue to flourish, the demand for generating personalized content is rapidly rising. Users increasingly seek to apply various editing methods to poli…
MVReward: Better Aligning and Evaluating Multi-View Diffusion Models with Human Preferences
Weitao Wang, Haoran Xu, Yuxiao Yang +3
Recent years have witnessed remarkable progress in 3D content generation. However, corresponding evaluation methods struggle to keep pace. Automatic approaches have proven challeng…