6 papers
PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation
Yuhang Huang, Xuan Lv, Junyan Xu +25
World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipula…
Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action
Yi Zhang, Yinda Chen, Che Liu +26
We present Pelican-Unify 1.0, the first embodied foundation model trained according to the principle of unification. Pelican-Unify 1.0 uses a single VLM as a unified understanding…
CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
Shilong Zou, Yuhang Huang, Renjiao Yi +2
We introduce a diffusion-based cross-domain image translator in the absence of paired training data. Unlike GAN-based methods, our approach integrates diffusion models to learn the…
AdaPower: Specializing World Foundation Models for Predictive Manipulation
Yuhang Huang, Shilong Zou, Jiazhao Zhang +3
World Foundation Models (WFMs) offer remarkable visual dynamics simulation capabilities, yet their application to precise robotic control remains limited by the gap between generat…
LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation
Yuhang Huang, Jiazhao Zhang, Shilong Zou +3
Predictive manipulation has recently gained considerable attention in the Embodied AI community due to its potential to improve robot policy performance by leveraging predicted sta…
Part-aware Shape Generation with Latent 3D Diffusion of Neural Voxel Fields
Yuhang Huang, SHilong Zou, Xinwang Liu +1
This paper presents a novel latent 3D diffusion model for the generation of neural voxel fields, aiming to achieve accurate part-aware structures. Compared to existing methods, the…