2 papers
cs.RO2026
Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline
Qing Yang, Xun Wang, Ziguan Wang +3
Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed…
cs.RO2026
MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation
Xingyuming Liu, Ruichun Ma, Heyu Guo +7
Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Language-Action (VLA) models for robotics rely solely on RGB observation…