From the 1 of 32 linked papers with an AI index.
2 citations · 2 across the 5 of their papers we have counts for
32 papers
Data Pyramid for Embodied Manipulation: A Survey
Yifan Ye, Yankai Fu, Yaoxu Lv +26
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…
Hierarchical Denoising For Multi-Step Visual Reasoning
Zezhong Qian, Xiaowei Chi, Chak-Wing Mak +9
The paper introduces HDR, a hierarchical denoising framework for causal video generation that enables multi-step visual reasoning with low-latency streaming, achieving higher succe…
WAM4D: Fast 4D World Action Model via Spatial Register Tokens
Ying Li, Xiaobao Wei, Jiajun Cao +10
World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video o…
PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing
Peng Li, Wangguandong Zheng, Yuan Liu +9
Detailed and photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, full-body reconstruction from a monocular RGB image r…
MV-WAM: Manifold-Aware World Action Model with Value Augmentation
Jintao Chen, Peidong Jia, Qingpo Wuwu +13
Achieving robust and generalizable manipulation across diverse environments remains a fundamental challenge in embodied robotics. Recent world action models achieve strong in-domai…
WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT
Zezhong Qian, Xiaowei Chi, Yu Qi +3
Recent World-Action (WA) models demonstrate strong generalization ability and data efficiency, but they typically rely on expert trajectories for training. This reliance limits the…