5 papers · 1 filter
LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction
Jin Lou, Zhiyuan Jing, Xupeng Wang +21
Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanato…
ViTaR: Visuo-Tactile Residual Adaptation for Foundation VLA Manipulation
Yi Wang, Renjun Wu, Jinyan Liu +1
As Vision-Language-Action (VLA) models scale toward real-world deployment, contact-rich manipulation exposes a critical blind spot: these policies encode broad visual-semantic prio…
GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation
Xiangyu Zhu, Renjun Wu, Luzhou Ge +2
Whole-body mobile manipulation requires coordinating mobile base and manipulator under shifting viewpoints, posing challenges in geometric perception and action generation. Current…
ReMAP-DP: Reprojected Multi-view Aligned PointMaps for Diffusion Policy
Xinzhang Yang, Renjun Wu, Jinyan Liu +1
Generalist robot policies built upon 2D visual representations excel at semantic reasoning but inherently lack the explicit 3D spatial awareness required for high-precision tasks.…
A Kung Fu Athlete Bot That Can Do It All Day: Highly Dynamic, Balance-Challenging Motion Dataset and Autonomous Fall-Resilient Tracking
Zhongxiang Lei, Lulu Cao, Xuyang Wang +3
Current humanoid motion tracking systems can execute routine and moderately dynamic behaviors, yet significant gaps remain near hardware performance limits and algorithmic robustne…