10 papers
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
Siyuan Xu, Tianshi Wang, Fengling Li +2
Vision-Language-Action models (VLAs) have demonstrated strong potential for embodied AI, yet their deployment on resource-limited robots remains challenging due to high memory and…
Non-Markovian Long-Horizon Robot Manipulation via Keyframe Chaining
Yipeng Chen, Wentao Tan, Lei Zhu +4
Existing Vision-Language-Action (VLA) models often struggle to generalize to long-horizon tasks due to their heavy reliance on immediate observations. While recent studies incorpor…
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
Chenyv Liu, Wentao Tan, Lei Zhu +4
Standard vision-language-action (VLA) models rely on fitting statistical data priors, limiting their robust understanding of underlying physical dynamics. Reinforcement learning en…
MOTIF: Learning Action Motifs for Few-shot Cross-Embodiment Transfer
Heng Zhi, Wentao Tan, Lei Zhu +4
While vision-language-action (VLA) models have advanced generalist robotic learning, cross-embodiment transfer remains challenging due to kinematic heterogeneity and the high cost…
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
Wenda Yu, Tianshi Wang, Fengling Li +2
Vision-Language-Action (VLA) models have demonstrated strong performance in robotic manipulation, yet their closed-loop deployment is hindered by the high latency and compute cost…
Generalizing Vision-Language Models with Dedicated Prompt Guidance
Xinyao Li, Yinjie Min, Hongbo Chen +3
Fine-tuning large pretrained vision-language models (VLMs) has emerged as a prevalent paradigm for downstream adaptation, yet it faces a critical trade-off between domain specifici…