1 paper
Boyang Shen, Kaixiang Yang, Hao Wang +4
Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal for action prediction. However, robotic…