20 papers
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
Xionghao Wu, Yijun Yang, Shiyang Zhou +17
Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect…
Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning
Yuxing Long, Lei Kang, Ziyan Yu +8
Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances, yet existing large models fall short, as no sufficiently diverse,…
PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion
Qingdong Xu, Jiajun Zhu, Shilin Zhu +4
We propose PatchScene, a novel diffusion-based framework for large-scale LiDAR scene completion. Unlike existing methods that rely on global latent representations or dense voxel g…
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
Jiyao Zhang, Mingxu Zhang, Yitong Peng +8
Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric bench…
HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning
Jiyao Zhang, Zimu Han, Junhan Wang +7
Robotic imitation learning faces a fundamental trade-off between modeling long-horizon dependencies and enabling fine-grained closed-loop control. Existing fixed-frequency action c…
FreeArtGS: Articulated Gaussian Splatting Under Free-moving Scenario
Hang Dai, Hongwei Fan, Han Zhang +3
The increasing demand for augmented reality and robotics is driving the need for articulated object reconstruction with high scalability. However, existing settings for reconstruct…