27 papers
DreamWAM: Beyond RGB Future Prediction for World Action Models
Shanglin Yuan, Weiheng Zhao, Xin Shi +6
World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-re…
Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models
Weiheng Zhao, Haoyi Jiang, Xin Shi +5
World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approaches face a fundamental dilemm…
IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
Zhengyu Zou, Hao Li, Kuixuan Jiao +7
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spa…
EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI
Xinjie Wang, Liu Liu, Taojun Ding +9
We present EmbodiedGen V2, a generative 3D world engine for building executable policy-ready environments for embodied intelligence. Sim-ready 3D asset generation has advanced rapi…
NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
Xiangyu Sun, Liu Liu, Seungkwon Yang +4
Pose-Free Feed-forward 3D Gaussian Splatting (3DGS) has recently emerged as a powerful paradigm for fast scene reconstruction. However, its performance degrades significantly in lo…
TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
Ziqian Wang, Yonghao He, Licheng Yang +6
Simulation provides a low-cost, scalable pathway to large-scale robotic manipulation data collection. However, existing 3D scene generation methods can rarely be applied directly t…