4 citations · 7 across the 7 of their papers we have counts for
11 papers · 1 filter
OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping
Zelong Lv, Sicheng Xu, Jianfeng Xiang +5
We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-consistent, and high-fidelity visual scenes. Given a single image, OctWo…
Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors
Hakyeong Kim, Ruicheng Wang, Chengtang Yao +2
Direct Time-of-Flight (dToF) sensors provide highly accurate metric depth and are more robust than indirect ToF systems in challenging real-world conditions. However, their high ma…
MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement
Lingyu Kong, Ruicheng Li, Ruicheng Wang +4
Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structu…
Map2World: Segment Map Conditioned Text to 3D World Generation
Jaeyoung Chung, Suyoung Lee, Jianfeng Xiang +2
3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising r…
Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data
Yizhao Xu, Hongyuan Zhu, Caiyun Liu +6
3D editing refers to the ability to apply local or global modifications to 3D assets. Effective 3D editing requires maintaining semantic consistency by performing localized changes…
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
Huizhi Liang, Yichao Shen, Yu Deng +5
Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D…