Showing cs.ROShow all
3 papers · 1 filter
cs.RO2026
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
Zixuan Wang, Yuxin Chen, Yuqi Liu +6
Vision-Language-Action (VLA) models typically map visual observations and linguistic instructions directly to control signals. This "black-box" mapping forces a single forward pass…
cs.RO2026
CLAW: Composable Language-Annotated Whole-body Motion Generation
Jianuo Cao, Yuxin Chen, Masayoshi Tomizuka
Training language-conditioned whole-body controllers for humanoid robots demands large-scale motion-language datasets. Existing approaches based on motion capture are costly and li…
cs.RO2026
StarVLA-: Reducing Complexity in Vision-Language-Action Systems
Jinhui Ye, Ning Gao, Senqiao Yang +7
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for building general-purpose robotic agents. However, the VLA landscape remains highly fragmented…