2 papers
cs.RO2026
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
Yongjie Bai, Zhouxia Wang, Yang Liu +8
Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlu…
cs.RO2025
RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model
Mingtong Dai, Lingbo Liu, Yongjie Bai +6
Vision-Language-Action (VLA) models have become a prominent paradigm for embodied intelligence, yet further performance improvements typically rely on scaling up training data and…