1 paper
Yongjie Bai, Zhouxia Wang, Yang Liu +8
Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlu…