8 papers
Teaching Tiny VLA Models Where to Look and How to Move
Lei Iok Tong, Iok Tong Lei, Qingchen Xie +6
Tiny Vision-Language-Action models are appealing for real-time robotic control, but reducing model scale often weakens two capabilities essential for manipulation: task-conditioned…
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning
Iok Tong Lei, QianZhi Li, Ying Jie Yap +5
Open-ended tabletop manipulation requires agents to not only understand natural language but also adapt to dynamic environments and execution failures. We present ACE (Agentic Cont…
OpenSPM: An Environment-Transferable Robotic Key Spatial Pose Memory and Closed-Loop High-Frequency Flow-Matching Action Generation Model
Iok Tong Lei, Qingchen Xie, Yifan Wang +2
Open-environment tabletop robotic manipulation requires systems to possess semantic understanding, precise geometric pose estimation, and high-frequency action generation. While en…
PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought
Ling Li, Bowen Liu, Zinuo Zhan +5
Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Trad…
VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection
Ling Li, Zhizhen Cai, Xinkun Wu +4
Grounding deictic gestures in natural images is fundamental to AR and human-robot collaboration, providing a basis for seamless spatial interaction. While Transformer-based visual…
From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper
Ling Li, Changjie Chen, Yuyan Wang +4
In multi-view 3D human pose estimation, models typically rely on images captured simultaneously from different camera views to predict a pose at a specific moment. While providing…