collaborators

6 papers

cs.RO2026

Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference

Zebin Yang, Qi Wang, Yunhe Wang +6

The paper introduces Jetson-PI, a system that enables real-time deployment of vision-language-action models on low-power edge devices like the Jetson Orin by using foresight-aligne…

cs.RO2026

Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation

Liang Heng, Jiadong Xu, Yiwen Wang +6

Relational object rearrangement (ROR) tasks (e.g., insert flower to vase) require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches e…

cs.RO2026

AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation

Juan Zhu, Zhanying Shao, Xiaoqi Li +4

Since current Vision-Language-Action (VLA) systems suffer from limited spatial perception and the absence of memory throughout manipulation, we investigate visual anchors as a mean…

cs.RO2025

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

Liang Heng, Xiaoqi Li, Shangqing Mao +9

Recent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection…

cs.RO2025

CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation

Xiaoqi Li, Lingyun Xu, Mingxu Zhang +8

In robotic, task goals can be conveyed through various modalities, such as language, goal images, and goal videos. However, natural language can be ambiguous, while images or video…

cs.CV2025

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

Xiaoqi Li, Jiaming Liu, Nuowei Han +4

The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide mode…