Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models
Jingzhou Luo, Yifan Wen, Yongjie Bai +3
Vision-Language-Action (VLA) models have shown strong performance on embodied manipulation, yet they remain brittle under visual observation changes, paraphrased language instructi…
cs.RO2026
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
Yongjie Bai, Zhouxia Wang, Yang Liu +8
Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlu…