9 papers
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
Changhua Xu, En Yu, Junyu Xuan +1
Vision--Language--Action (VLA) models bridge multimodal reasoning with physical control, but adapting them to new tasks with scarce demonstrations remains unreliable. While fine-tu…
NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria
Eason Yu, Tzu Hao Liu, Clément L. Canonne +4
Finding Nash equilibria in two-player zero-sum imperfect-information games remains a central challenge in multi-agent reinforcement learning. Recent multi-round regularization meth…
Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
Zhe Li, Cheng Chi, Yangyang Wei +9
Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion fro…
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
Zhe Li, Cheng Chi, Boan Zhu +11
Humans learn locomotion through visual observation, interpreting visual content first before imitating actions. However, state-of-the-art humanoid locomotion systems rely on either…
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
Shuanghao Bai, Wenxuan Song, Jiayi Chen +14
Recent advances in vision, language, and multimodal learning have significantly accelerated progress in robotic foundation models, with robotic manipulation remaining one of the mo…
From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
Zhe Li, Cheng Chi, Yangyang Wei +7
Natural language offers a natural interface for humanoid robots, but existing language-guided humanoid locomotion pipelines remain cumbersome and untrustworthy. They typically deco…