From the 1 of 29 linked papers with an AI index.
29 papers
DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving
Zebin Xing, Yupeng Zheng, Qiang Chen +10
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across perception, language, and p…
VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
Yupeng Zheng, Kai Zou, Bin Liu +1
VisCo introduces a training-efficient self-compression framework that reuses a pretrained vision-language model as an intrinsic autoencoder to compress visual tokens into a small s…
VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation
Shuai Tian, Yupeng Zheng, Yuhang Zheng +7
Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observat…
TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation
Yujie Zang, Yuhang Zheng, Xian Nie +7
Contact-rich manipulation requires robots to continuously perceive and regulate evolving physical interactions under dynamic contact transitions or complex surface geometries. Rece…
Learning High-Frequency Continuous Action Chunks in Latent Space
Kunyun Wang, Yuhang Zheng, Yupeng Zheng +2
Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal consistency at moderate action…
Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning
Yuan Liu, Haoran Li, Shuai Tian +5
Pretrained on large-scale and diverse datasets, VLA models demonstrate strong generalization and adaptability as general-purpose robotic policies. However, Supervised Fine-Tuning (…