1 paper · 1 filter
Jinyue Bian, Zhaoxing Zhang, Zhengyu Liang +5
The Visual-Language-Action (VLA) models can follow text instructions according to visual observations of the surrounding environment. This ability to map multimodal inputs to actio…