1 paper
Zhide Zhong, Haodong Yan, Junfeng Li +8
Many Vision-Language-Action (VLA) models are built upon an internal world model trained via next-frame prediction ``vt→vt+1''. However, this paradigm attempts to…