1 paper
Yupeng Zheng, Xiang Li, Songen Gu +11
Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction obje…