Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding
Hung T. Ho, Khanh D. Nguyen, Quang D. Nguyen +5
Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework…
cs.RO2026
vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen +5
Deploying vision--language--action (VLA) models on robots requires adapting model-specific inference pipelines to heterogeneous processors and limited onboard memory. We present vl…