activity
20242026
most citedDriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

4 citations · 5 across the 23 of their papers we have counts for

collaborators

24 papers

cs.RO2026

Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment

Boyuan Wang, Zhenyuan Zhang, Zhiqin Yang +16

Autonomous data collection governs the volume and quality of real-world trajectories for manipulation policy learning. Existing pipelines reduce human effort via self-resetting, VL…

cs.RO2026

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

GigaWorld Team, Angen Ye, Angyuan Ma +26

World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physicall…

cs.RO2026

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

Angen Ye, Weijie Ke, Xiaofeng Wang +7

World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynami…

cs.RO2026

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

GigaWorld Team, Angyuan Ma, Boyuan Wang +24

Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow,…

cs.CV2026

ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

Boyuan Wang, Xiaofeng Wang, Yongkang Li +9

Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for per-scene optimization, recov…

cs.CV2026

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

Yang Zhou, Xiaofeng Wang, Hao Shao +8

Recently, world-action models (WAM) have emerged to bridge vision-language-action (VLA) models and world models, unifying their reasoning and instruction-following capabilities and…