collaborators

6 papers

cs.RO2025

PoseDiff: A Unified Diffusion Model Bridging Robot Pose Estimation and Video-to-Action Control

Haozhuo Zhang, Michele Caprio, Jing Shao +4

We present PoseDiff, a conditional diffusion model that unifies robot state estimation and control within a single framework. At its core, PoseDiff maps raw visual observations int…

cs.RO2025

Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots

Wei Cui, Haoyu Wang, Wenkang Qin +19

Humanoid robot technology is advancing rapidly, with manufacturers introducing diverse heterogeneous visual perception modules tailored to specific scenarios. Among various percept…

cs.CV2025

Occupancy World Model for Robots

Zhang Zhang, Qiang Zhang, Wei Cui +12

Understanding and forecasting the scene evolutions deeply affect the exploration and decision of embodied agents. While traditional methods simulate scene evolutions through trajec…

cs.RO2025

RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots

Zhang Zhang, Qiang Zhang, Wei Cui +7

3D occupancy prediction enables the robots to obtain spatial fine-grained geometry and semantics of the surrounding scene, and has become an essential task for embodied perception.…

cs.RO2025

EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

Yi Zhang, Qiang Zhang, Xiaozhu Ju +13

While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex l…

cs.RO2025

HumanoidPano: Hybrid Spherical Panoramic-LiDAR Cross-Modal Perception for Humanoid Robots

Qiang Zhang, Zhang Zhang, Wei Cui +13

The perceptual system design for humanoid robots poses unique challenges due to inherent structural constraints that cause severe self-occlusion and limited field-of-view (FOV). We…