10 papers · 1 filter
DreamWAM: Beyond RGB Future Prediction for World Action Models
Shanglin Yuan, Weiheng Zhao, Xin Shi +6
World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-re…
EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI
Xinjie Wang, Liu Liu, Taojun Ding +9
We present EmbodiedGen V2, a generative 3D world engine for building executable policy-ready environments for embodied intelligence. Sim-ready 3D asset generation has advanced rapi…
Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors
Youzhi Liu, Li Gao, Yifei Qian +3
Quadruped robots have achieved remarkable locomotion, yet their behavioral repertoire remains confined to a few gaits--far from the expressive, companion-like presence long envisio…
GeoFlow-SLAM++: A Robust Multi-Camera Visual-Inertial SLAM System with Relocalization
Wei Feng, Tingyang Xiao, Liu Liu +2
Monocular and RGB-D visual-inertial SLAM systems remain susceptible to limited field of view, sensor-specific failure modes, and unreliable cross-session relocalization. To address…
HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory
Xiaolin Zhou, Liu Liu, Tingyang Xiao +9
LLM agents follow a practical execution loop in digital environments: they reason over structured states, invoke tools, inspect feedback, and revise actions. Extending this loop to…
Monocular 3D Occupancy Perception for Robots on Sidewalks via Hybrid 2D-3D Learning
Yukai Ma, Joe Lin, Liu Liu +5
Sidewalks in the real world are crowded, cluttered, and less structured than roads, making 3D occupancy prediction a key ingredient for the safe navigation of mobile robots such as…