1 citations · 2 across the 18 of their papers we have counts for
4 papers · 1 filter
Learning Foresight without Explicit Trajectories for 3D Diffusion Policies
Zhongbo Zhang, Zaibin Zhang, Yifan Wang +3
3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasib…
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
Zhongbo Zhang, Jiayi Jin, Yifan Wang +4
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spati…
LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models
Lin Liu, Zhicheng Bao, Lu Zhang +7
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100…
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
Zaibin Zhang, Junlan Xiao, Zhongbo Zhang +11
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most rep…