activity
20242026
collaborators

22 papers

cs.RO2026

DreamWAM: Beyond RGB Future Prediction for World Action Models

Shanglin Yuan, Weiheng Zhao, Xin Shi +6

World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-re…

cs.CV2026

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Weiheng Zhao, Haoyi Jiang, Xin Shi +5

World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approaches face a fundamental dilemm…

cs.CV2026

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

Kangsheng Duan, Ziyang Xu, Wenyu Liu +3

While 10B-level industrial foundation models have pushed the boundaries of image inpainting, their prohibitive computational costs severely hinder practical deployment. Constructin…

cs.RO2026

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

Shanglin Yuan, Weiheng Zhao, Xianda Guo +4

Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiot…

cs.CV2026

Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning

Yu Zhu, Yongkang Li, Wenjie Zhu +5

Recent studies have explored Vision-Language Models (VLMs) for food analysis. However, most existing methods rely primarily on supervised fine-tuning (SFT), which often limits reas…

cs.CV2026

UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving

Yongkang Li, Lijun Zhou, Sixu Yan +11

Vision-Language-Action (VLA) models have recently emerged in autonomous driving, with the promise of leveraging rich world knowledge to improve the cognitive capabilities of drivin…