4 papers
Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
Yifan Ye, Jiaqi Ma, Jun Cen +1
Vision-Language-Action (VLA) models pretrained on large-scale multimodal datasets have emerged as powerful foundations for robotic perception and control. However, their massive sc…
Temporal Realism Evaluation of Generated Videos Using Compressed-Domain Motion Vectors
Mert Onur Cakiroglu, Idil Bilge Altun, Zhihe Lu +2
Temporal realism remains a central weakness of current generative video models, as most evaluation metrics prioritize spatial appearance and offer limited sensitivity to motion. We…
Self-evolved Imitation Learning in Simulated World
Yifan Ye, Jun Cen, Jing Chen +1
Imitation learning has been a trend recently, yet training a generalist agent across multiple tasks still requires large-scale expert demonstrations, which are costly and labor-int…
Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion
Yuguang Zhang, Kuangpu Guo, Zhihe Lu +2
Federated learning (FL) enables collaborative model training across decentralized clients without sharing local data, but is challenged by heterogeneity in data, computation, and c…