5 papers
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
Siyuan Yang, Yang Zhang, Haoran He +4
Vision-Language-Action (VLA) models, trained via flow-matching or diffusion objectives, excel at learning complex behaviors from large-scale, multi-modal datasets (e.g., human tele…
Directed-MAML: Meta Reinforcement Learning Algorithm with Task-directed Approximation
Yang Zhang, Huiwen Yan, Mushuang Liu
Model-Agnostic Meta-Learning (MAML) is a versatile meta-learning framework applicable to both supervised learning and reinforcement learning (RL). However, applying MAML to meta-re…
Eyes Will Shut: A Vision-Based Next GPS Location Prediction Model by Reinforcement Learning from Visual Map Feed Back
Ruixing Zhang, Yang Zhang, Tongyu Zhu +2
Next Location Prediction is a fundamental task in the study of human mobility, with wide-ranging applications in transportation planning, urban governance, and epidemic forecasting…
FABG : End-to-end Imitation Learning for Embodied Affective Human-Robot Interaction
Yanghai Zhang, Changyi Liu, Keting Fu +3
This paper proposes FABG (Facial Affective Behavior Generation), an end-to-end imitation learning system for human-robot interaction, designed to generate natural and fluid facial…
Pre-Trained Video Generative Models as World Simulators
Haoran He, Yang Zhang, Liang Lin +2
Video generative models pre-trained on large-scale internet datasets have achieved remarkable success, excelling at producing realistic synthetic videos. However, they often genera…