1 paper
Jianke Zhang, Yuanfei Luo, Yucheng Hu +6
Vision--language--action (VLA) models are typically built by fine-tuning a pretrained vision--language model (VLM) on action data. However, we show that this standard recipe system…