6 papers
PoseDiff: A Unified Diffusion Model Bridging Robot Pose Estimation and Video-to-Action Control
Haozhuo Zhang, Michele Caprio, Jing Shao +4
We present PoseDiff, a conditional diffusion model that unifies robot state estimation and control within a single framework. At its core, PoseDiff maps raw visual observations int…
Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
Wei Cui, Haoyu Wang, Wenkang Qin +19
Humanoid robot technology is advancing rapidly, with manufacturers introducing diverse heterogeneous visual perception modules tailored to specific scenarios. Among various percept…
Occupancy World Model for Robots
Zhang Zhang, Qiang Zhang, Wei Cui +12
Understanding and forecasting the scene evolutions deeply affect the exploration and decision of embodied agents. While traditional methods simulate scene evolutions through trajec…
RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots
Zhang Zhang, Qiang Zhang, Wei Cui +7
3D occupancy prediction enables the robots to obtain spatial fine-grained geometry and semantics of the surrounding scene, and has become an essential task for embodied perception.…
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
Yi Zhang, Qiang Zhang, Xiaozhu Ju +13
While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex l…
HumanoidPano: Hybrid Spherical Panoramic-LiDAR Cross-Modal Perception for Humanoid Robots
Qiang Zhang, Zhang Zhang, Wei Cui +13
The perceptual system design for humanoid robots poses unique challenges due to inherent structural constraints that cause severe self-occlusion and limited field-of-view (FOV). We…