8 papers
Data Pyramid for Embodied Manipulation: A Survey
Yifan Ye, Yankai Fu, Yaoxu Lv +26
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…
Auditing Generalization in AI-Generated Video Detection: A Six-Control Protocol and the VidAudit Toolkit
Mert Onur Cakiroglu, Zhihe Lu, Mehmet Dalkilic +1
AI-generated video detection benchmarks such as GenVidBench and AIGVDBench are the de facto leaderboards, yet most evaluation protocols leave uncontrolled confounds that can inflat…
Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination
Jiajun Li, Tiecheng Guo, Yifan Ye +9
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely o…
Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation
Yunfan Lou, Yifan Ye, Yankai Fu +7
World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on v…
Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
Yifan Ye, Jiaqi Ma, Jun Cen +1
Vision-Language-Action (VLA) models pretrained on large-scale multimodal datasets have emerged as powerful foundations for robotic perception and control. However, their massive sc…
Temporal Realism Evaluation of Generated Videos Using Compressed-Domain Motion Vectors
Mert Onur Cakiroglu, Idil Bilge Altun, Zhihe Lu +2
Temporal realism remains a central weakness of current generative video models, as most evaluation metrics prioritize spatial appearance and offer limited sensitivity to motion. We…