works on

From the 1 of 21 linked papers with an AI index.

activity
20242026
collaborators

21 papers

cs.RO2026

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

Peize Li, Ruimeng Zhang, Ru Zhang +3

Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real…

cs.RO2026

Data Pyramid for Embodied Manipulation: A Survey

Yifan Ye, Yankai Fu, Yaoxu Lv +26

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…

cs.CV2026

Hierarchical Denoising For Multi-Step Visual Reasoning

Zezhong Qian, Xiaowei Chi, Chak-Wing Mak +9

The paper introduces HDR, a hierarchical denoising framework for causal video generation that enables multi-step visual reasoning with low-latency streaming, achieving higher succe…

cs.CV2026

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

Ying Li, Xiaobao Wei, Jiajun Cao +10

World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video o…

cs.RO2026

LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes

Haozhuo Zhang, Jingkai Sun, Michele Caprio +5

Physics-based human motion control can make a simulated character walk, sit, and manipulate objects with high physical realism. Almost always, though, this happens in short, isolat…

cs.RO2026

LaST: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

Zhuoyang Liu, Jiaming Liu, Hao Chen +11

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future obs…