collaborators

8 papers

cs.RO2026

Data Pyramid for Embodied Manipulation: A Survey

Yifan Ye, Yankai Fu, Yaoxu Lv +26

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…

cs.CV2026

Auditing Generalization in AI-Generated Video Detection: A Six-Control Protocol and the VidAudit Toolkit

Mert Onur Cakiroglu, Zhihe Lu, Mehmet Dalkilic +1

AI-generated video detection benchmarks such as GenVidBench and AIGVDBench are the de facto leaderboards, yet most evaluation protocols leave uncontrolled confounds that can inflat…

cs.RO2026

Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

Jiajun Li, Tiecheng Guo, Yifan Ye +9

World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely o…

cs.RO2026

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

Yunfan Lou, Yifan Ye, Yankai Fu +7

World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on v…

cs.RO2025

Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models

Yifan Ye, Jiaqi Ma, Jun Cen +1

Vision-Language-Action (VLA) models pretrained on large-scale multimodal datasets have emerged as powerful foundations for robotic perception and control. However, their massive sc…

cs.CV2025

Temporal Realism Evaluation of Generated Videos Using Compressed-Domain Motion Vectors

Mert Onur Cakiroglu, Idil Bilge Altun, Zhihe Lu +2

Temporal realism remains a central weakness of current generative video models, as most evaluation metrics prioritize spatial appearance and offer limited sensitivity to motion. We…