collaborators

5 papers

cs.RO2026

DreamWAM: Beyond RGB Future Prediction for World Action Models

Shanglin Yuan, Weiheng Zhao, Xin Shi +6

World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-re…

cs.CV2026

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Weiheng Zhao, Haoyi Jiang, Xin Shi +5

World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approaches face a fundamental dilemm…

cs.RO2026

TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors

Pinhan Fu, Xianda Guo, Xuetao Li +5

Vision-Language-Action (VLA) models are deployed through pipelines that end users cannot audit, and a poisoned VLA can behave normally on clean observations while a small visual tr…

cs.RO2026

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

Shanglin Yuan, Weiheng Zhao, Xianda Guo +4

Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiot…

cs.CV2025

SuperCLIP: CLIP with Simple Classification Supervision

Weiheng Zhao, Zilong Huang, Jiashi Feng +1

Contrastive Language-Image Pretraining (CLIP) achieves strong generalization in vision-language tasks by aligning images and texts in a shared embedding space. However, recent find…