2 papers
cs.LG2026
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
Sen Wang, R. Gnana Praveen, Bidhan Roy +1
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transf…
cs.RO2026
WSA: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control
Jiahao Jiang, Jianing Zhang, Zhenhan Yin +8
Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for generalist robotic systems to date. By leveraging imitation learning on…