8 papers
How Should Vision-Language-Action Models Use Proprioceptive State?
Yiren Zhao, Ziyang Chen, Ziyang Rao +5
Recent Vision-Language-Action (VLA) models almost universally take robot proprioceptive state as input, yet wire it in incompatible ways -- serialized into text prompts, projected…
Source-Lifted Flow Matching for Intervenable Multimodal Imitation
He Zhang, Ying Sun, Pengteng Li +6
Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stochasticity is largely passive: repeated sa…
PHASER: Phase-Aware and Semantic Experience Replay for Vision-Language-Action Models
Ziyang Chen, Shaoguang Wang, Weiyu Guo +5
Vision-Language-Action (VLA) models have achieved remarkable success in language-conditioned robotic manipulation. However, deploying these models in open-ended environments requir…
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
Shaoguang Wang, Weiyu Guo, Ziyang Chen +2
Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive computational cost of processing dense frame sequences.…
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
Jianxiang He, Meisheng Hong, Jungang Li +3
Multimodal large language models (MLLMs) demonstrate exceptional performance in vision-language tasks, yet their processing of long videos is constrained by input context length an…
A Brain-inspired Embodied Intelligence for Fluid and Fast Reflexive Robotics Control
Weiyu Guo, He Zhang, Pengteng Li +7
Recent advances in embodied intelligence have leveraged massive scaling of data and model parameters to master natural-language command following and multi-task control. In contras…