8 papers
FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
Ruicheng Li, Qixiu Li, Ruichun Ma +8
Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by condi…
MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement
Lingyu Kong, Ruicheng Li, Ruicheng Wang +4
Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structu…
Map2World: Segment Map Conditioned Text to 3D World Generation
Jaeyoung Chung, Suyoung Lee, Jianfeng Xiang +2
3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising r…
NeAR: Coupled Neural Asset-Renderer Stack
Hong Li, Chongjie Ye, Houyuan Chen +12
Neural asset authoring and neural rendering have traditionally evolved as disjoint paradigms: one generates digital assets for fixed graphics pipelines, while the other maps conven…
Native and Compact Structured Latents for 3D Generation
Jianfeng Xiang, Xiaoxue Chen, Sicheng Xu +8
Recent advancements in 3D generative modeling have significantly improved the generation realism, yet the field is still hampered by existing representations, which struggle to cap…
MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
Ruicheng Wang, Sicheng Xu, Yue Dong +6
We propose MoGe-2, an advanced open-domain geometry estimation model that recovers a metric scale 3D point map of a scene from a single image. Our method builds upon the recent mon…