2 papers
cs.RO2026
Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification
Chunyu Qi, Zhuoran Song, Jian Weng +6
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hin…
cs.AI2026
A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference
Zhuoran Song, Haozhe Jiang, Chunyu Qi +4
Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators,…