2 papers
cs.CV2026
MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision
Zijie Zhu, Weiren Cai, Yizhou Wang +4
Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity understanding, robot learning, and augmented reality. Existing systems…
cs.RO2026
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
Yuran Wang, Siqiao Huang, Mingleyang Li +21
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are…