2 papers
cs.RO2026
LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction
Jin Lou, Zhiyuan Jing, Xupeng Wang +21
Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanato…
cs.RO2026
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
Junhao Cai, Zetao Cai, Jiafei Cao +39
Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, b…