11 papers
DreamWAM: Beyond RGB Future Prediction for World Action Models
Shanglin Yuan, Weiheng Zhao, Xin Shi +6
World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-re…
Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models
Weiheng Zhao, Haoyi Jiang, Xin Shi +5
World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approaches face a fundamental dilemm…
MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model
Shanglin Yuan, Weiheng Zhao, Xianda Guo +4
Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiot…
3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
Ze-Xin Yin, Liu Liu, Xinjie Wang +4
Compositional 3D scene generation from a single view requires the simultaneous recovery of scene layout and 3D assets. Existing approaches mainly fall into two categories: feed-for…
DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation
Ze-Xin Yin, Jiaxiong Qiu, Liu Liu +5
The labor- and experience-intensive creation of 3D assets with physically based rendering (PBR) materials demands an autonomous 3D asset creation pipeline. However, most existing 3…
IRIS-SLAM: Unified Geo-Instance Representations for Robust Semantic Localization and Mapping
Tingyang Xiao, Liu Liu, Wei Feng +6
Geometry foundation models have significantly advanced dense geometric SLAM, yet existing systems often lack deep semantic understanding and robust loop closure capabilities. Meanw…