1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Bryce Grant, Xijia Zhao, Peng Wang
Vision-Language-Action (VLA) models combine perception, language, and motor control in a single architecture, yet how they translate multimodal inputs into actions remains poorly u…