collaborators

8 papers

cs.RO2026

How Should Vision-Language-Action Models Use Proprioceptive State?

Yiren Zhao, Ziyang Chen, Ziyang Rao +5

Recent Vision-Language-Action (VLA) models almost universally take robot proprioceptive state as input, yet wire it in incompatible ways -- serialized into text prompts, projected…

cs.RO2026

Source-Lifted Flow Matching for Intervenable Multimodal Imitation

He Zhang, Ying Sun, Pengteng Li +6

Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stochasticity is largely passive: repeated sa…

cs.RO2026

PHASER: Phase-Aware and Semantic Experience Replay for Vision-Language-Action Models

Ziyang Chen, Shaoguang Wang, Weiyu Guo +5

Vision-Language-Action (VLA) models have achieved remarkable success in language-conditioned robotic manipulation. However, deploying these models in open-ended environments requir…

cs.CV2026

Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding

Shaoguang Wang, Weiyu Guo, Ziyang Chen +2

Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive computational cost of processing dense frame sequences.…

cs.CV2026

VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding

Jianxiang He, Meisheng Hong, Jungang Li +3

Multimodal large language models (MLLMs) demonstrate exceptional performance in vision-language tasks, yet their processing of long videos is constrained by input context length an…

cs.RO2026

A Brain-inspired Embodied Intelligence for Fluid and Fast Reflexive Robotics Control

Weiyu Guo, He Zhang, Pengteng Li +7

Recent advances in embodied intelligence have leveraged massive scaling of data and model parameters to master natural-language command following and multi-task control. In contras…