4 papers
vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen +5
Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware on which robots actually run.…
Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
Khanh-Binh Nguyen, Phuoc-Nguyen Bui, Hyunseung Choo +1
Vision-language models (VLMs) exhibit remarkable zero-shot generalization but suffer performance degradation under distribution shifts in downstream tasks, particularly in the abse…
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
Phuoc-Nguyen Bui, Khanh-Binh Nguyen, Hyunseung Choo
Contrastive vision-language models excel in zero-shot image recognition but face challenges in few-shot scenarios due to computationally intensive offline fine-tuning using prompt…
Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models
Phuoc-Nguyen Bui, Khanh-Binh Nguyen, Hyunseung Choo
Vision-language models (VLMs) like CLIP excel in zero-shot learning but often require resource-intensive training to adapt to new tasks. Prompt learning techniques, such as CoOp an…