5 papers · 1 filter
WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression
Maeve Zhang, Rain Sun, Xiang Wang +22
Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and ro…
HOST:Robots Acquire Manipulation Skills in Seconds from a Single Human Video
Guangyan Chen, Meiling Wang, Te Cui +9
The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-…
WALL-WM: Carving World Action Modeling at the Event Joints
Shalfun Li, Victor Yao, Charles Yang +29
WALL-WM is a World Action Model that shifts video-action learning from chunk-centric optimization to event-grounded Vision-Language-Action pretraining, using semantically coherent…
Wall-OSS-0.5 Technical Report
Ryan Yu, Pushi Zhang, Starrick Liu +24
Large-scale Vision-Language-Action (VLA) pretraining is increasingly adopted as the foundation for robot policies, yet the evidence for pretrained VLAs is almost invariably reporte…
Igniting VLMs toward the Embodied Space
Andy Zhai, Brae Liu, Bruno Fang +17
While foundation models show remarkable progress in language and vision, existing vision-language models (VLMs) still have limited spatial and embodiment understanding. Transferrin…