24 papers
Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
Houze Xu, Jizhong Li, Ziyi Ye
Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still fac…
Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation
Shengqi Xu, Guojin Zhong, Yang Liu +7
Visuo-Tactile policies leveraging optical tactile sensors have shown great promise in contact-rich manipulation. These sensors achieve high spatial resolution and multi-dimensional…
ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation
Tianyi Lu, Hui Zhang, Zijie Diao +8
Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiting their capacity for reasoning-intensive long-horizon tasks. To add…
Teach Multimodal Recommendation Model to See via Personalized Visual Extraction and Adaptive Learning
Yutong Li, Xinyi Zhang, Ziyi Ye +2
Multimodal sequential recommendation (MSR) incorporates textual and visual information to improve recommendation quality. However, recent studies and our empirical analysis show th…
ActiveMimic: Egocentric Video Pretraining with Active Perception
Xingyao Lin, Guojin Zhong, Tianyi Lu +4
Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform those pretrained on robot data. We…
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
Shengyu Si, Yuanzhuo Lu, Ruimeng Yang +3
Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate tra…