6 papers
Learning Foresight without Explicit Trajectories for 3D Diffusion Policies
Zhongbo Zhang, Zaibin Zhang, Yifan Wang +3
3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasib…
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
Zhongbo Zhang, Jiayi Jin, Yifan Wang +4
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spati…
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
Changbo Yan, Zhongbo Zhang, Zaibin Zhang +3
3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may be partially occluded or tightly intermingled with visually…
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
Zaibin Zhang, Junlan Xiao, Zhongbo Zhang +11
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most rep…
Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning
Tieliang Gong, Zhongbo Zhang, Wen Wen +1
Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies…
Think3D: Thinking with Space for Spatial Reasoning
Zaibin Zhang, Yuhan Wu, Lianjie Jia +10
While Vision-Language Models (VLMs) excel at 2D visual understanding, they remain constrained by 2D-centric paradigm that severely limits genuine 3D spatial reasoning. To bridge th…