2 papers
cs.RO2026
OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
Yushan Liu, Peibo Sun, Shoujie Li +7
World Action Models (WAMs) enhance Vision-Language-Action policies by jointly predicting scene evolution and robot actions, but existing methods usually represent the predicted wor…
cs.CV2025
Perception Encoder: The best visual embeddings are not at the output of the network
Daniel Bolya, Po-Yao Huang, Peize Sun +15
We introduce Perception Encoder (PE), a state-of-the-art vision encoder for image and video understanding trained via simple vision-language learning. Traditionally, vision encoder…