collaborators

7 papers

cs.CV2026

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Yuyang Zhang, Wenyao Zhang, Zekun Qi +7

World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense mult…

cs.RO2026

VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

Jingwen Sun, Wenyao Zhang, Zekun Qi +6

Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pi…

cs.RO2026

ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models

Baorui Peng, Wenyao Zhang, Liang Xu +5

Recently, video-based world models that learn to simulate the dynamics have gained increasing attention in robot learning. However, current approaches primarily emphasize visual ge…

cs.CV2025

Reasoning in Space via Grounding in the World

Yiming Chen, Zekun Qi, Wenyao Zhang +3

In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the Grounded-Spatial Reasoner (GS-Reasoner) to explore the effective spatial…

cs.CV2025

Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation

Wenyao Zhang, Hongsi Liu, Bohan Li +7

Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this cha…

cs.RO2025

SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation

Zekun Qi, Wenyao Zhang, Yufei Ding +15

While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation-a key factor in 6-DoF fine-grained manipulation. Traditional p…