4 papers
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
Haoxiang Ma, Junhao Cai, Xiaoxu Xu +26
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In pract…
WSA: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control
Jiahao Jiang, Jianing Zhang, Zhenhan Yin +8
Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for generalist robotic systems to date. By leveraging imitation learning on…
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
Zhenhan Yin, Xuanhan Wang, Jiahao Jiang +8
While leveraging abundant human videos and simulated robot data poses a scalable solution to the scarcity of real-world robot data, the generalization capability of existing vision…
Pseudo-Label Refinement for Robust Wheat Head Segmentation via Two-Stage Hybrid Training
Jiahao Jiang, Zhangrui Yang, Xuanhan Wang +1
This extended abstract details our solution for the Global Wheat Full Semantic Segmentation Competition. We developed a systematic self-training framework. This framework combines…