4 citations · 4 across the 10 of their papers we have counts for
10 papers
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
Ruiteng Zhao, Zhengshen Zhang, Yue Su +6
World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both ali…
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
Yuhao Pan, Haosong Peng, Zhengshen Zhang +8
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-graine…
World Guidance: World Modeling in Condition Space for Action Generation
Yue Su, Sijin Chen, Haixin Shi +7
Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, e…
3D Affordance Keypoint Detection for Robotic Manipulation
Zhiyang Liu, Ruiteng Zhao, Lei Zhou +6
This paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts' functionality. The propo…
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
Haosong Peng, Hao Li, Yalun Dai +8
General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., cam…
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
Zhengshen Zhang, Hao Li, Yalun Dai +10
Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptabilit…